Structured JSON AI Prompt for Realistic Smartphone Flash Street Fashion Photography

A structured JSON AI prompt guide for generating realistic night street fashion portraits with authentic smartphone flash lighting, 26mm optics, skin texture, a

tau · October 5, 2026

#AIImages #PromptEngineering #ChatGPT #SmartphonePhotography #StreetFashion #Tips

Structured JSON AI Prompt for Realistic Smartphone Flash Street Fashion Photography

AI image creator and prompt engineer Hana (@jiwooteasing) has published a structured JSON prompt architecture designed to produce hyperrealistic night street fashion photography, capturing the raw, candid aesthetic of direct smartphone flash lighting.

Realistic Y2K street fashion portrait captured at night on an urban street with direct smartphone flash lighting

Image source: @jiwooteasing / X

Conventional, unstructured natural-language prompts frequently suffer from prompt bleeding and attention decay. In traditional image generation workflows, blending lighting instructions, subject anatomy, garment styling, and camera optics into a single comma-separated paragraph often produces an artificial plastic sheen (commonly dubbed 'AI slop') or defaults to unmotivated studio rim lighting. By decomposing every visual dimension into explicit, modular key-value pairs within a JSON specification, this workflow guides foundation models to faithfully render the optical characteristics of a 26mm lens, steep direct flash falloff, delicate skin pores with specular oil sheen, and realistic Y2K street fashion textures.

Architectural Advantages of Structured JSON Prompting

Structuring generative image prompts into formalized JSON objects delivers several distinct engineering benefits over freeform descriptive prose:

  • Eliminating Attribute Bleeding: Segregating wardrobe attributes (outfit) from environmental parameters (scene) ensures that pastel pink knitwear and medium-wash denim tones do not erroneously bleed into background dark-gray roller shutters or ambient shadow tones.
  • Enforcing Optical and Lighting Priority: When technical camera parameters (camera) and flash dynamics (lighting) are positioned as distinct top-level keys rather than secondary clauses buried behind physical character descriptions, the model's text encoder treats optical fidelity with equal priority during visual synthesis.
  • Modular Reusability and Lookbook Scalability: Instead of reconstructing an entire prompt from scratch for each shot, creators can swap isolated nested values—such as updating outfit, adjusting pose, or relocating the scene—to generate an entire coherent fashion lookbook with uniform stylistic and lighting continuity.

Complete Night Street Smartphone Flash Fashion JSON Prompt

The complete prompt specification below can be fed directly into ChatGPT or any foundation image generation pipeline capable of parsing structured text blocks:

{
  "direction": "A realistic smartphone lifestyle fashion photo captured outdoors at night on a quiet urban street. A beautiful adult South Korean woman in her early 20s stands naturally in front of closed industrial roller shutters illuminated by a single smartphone flash. Candid, stylish, and effortless rather than professionally staged.",

  "mood": "Quiet city night, youthful, casual, confident, spontaneous, slightly nostalgic Y2K street fashion.",

  "model": "Beautiful adult South Korean woman in her early 20s, approximately 163 cm tall, with elegant Korean proportions, naturally slim feminine physique, naturally full proportional bust with realistic softness and weight, slim toned waist, subtle hourglass silhouette, softly rounded hips, graceful shoulders, defined collarbones, and long slender legs.",

  "body": "Elegant Korean proportions with naturally feminine curves, realistic bust volume and gravity, slim ribcage, visible collarbones, toned waist, subtle hourglass silhouette, softly rounded hips, graceful shoulders, and long slender legs.",

  "skin": "Warm fair porcelain skin with subtle peach undertones, fine realistic texture, gentle pores, soft natural sheen, and realistic smartphone-flash reflections across the face, shoulders, collarbones, arms, and abdomen.",

  "outfit": {
    "top": "Fitted pastel pink ribbed camisole with a straight neckline and sparkling rhinestone shoulder straps. Natural stretch and believable contour.",
    "bottom": "Relaxed straight-leg medium-wash blue jeans worn at an extremely ultra-low-rise position, naturally resting on the lowest part of the hip bones while remaining securely worn. The low waistband exposes the lower abdomen, hip bones, waistline, and subtle V-line without revealing underwear. Simple black leather belt.",
    "accessories": "Minimal silver necklace with a tiny pendant. No additional accessories."
  },

  "scene": "Quiet urban street at night with large dark-gray industrial roller shutters, concrete pavement, and almost no traffic. Dark surroundings allow the smartphone flash to naturally isolate the subject. No colorful neon lights or busy distractions.",

  "pose": "Standing facing the camera with both arms naturally lifted behind the head, gently holding the back of the hair. Elbows point outward, shoulders remain relaxed, torso faces forward, and one hip shifts subtly to create a natural asymmetrical posture.",

  "composition": "Vertical 4:5 smartphone composition, photographed from approximately 1.8 meters away. Framed from head to upper thighs. Subject occupies around 80% of the frame while preserving enough dark street environment for context.",

  "lighting": "Single direct smartphone flash close to the camera, creating bright frontal illumination with natural falloff. Background remains dark while the subject is evenly lit, with realistic flash reflections on skin, hair, and clothing.",

  "camera": "Modern smartphone rear camera, approximately 26mm equivalent lens. Authentic smartphone flash photography, natural HDR, handheld framing, subtle JPEG compression, slight digital noise in darker areas, and genuine smartphone color science.",

  "style": "Photorealistic smartphone lifestyle fashion photography, authentic late-night social-media aesthetic, realistic anatomy, believable clothing behavior, natural flash lighting, detailed skin texture, and genuine handheld quality.",

  "negative": "long glamorous hair, plastic skin, exaggerated anatomy, bad hands, extra fingers, distorted limbs, CGI, anime, studio lighting, excessive bokeh, over-retouched skin, blurry face, text, watermark, logo"
}

Four Critical Levers for Direct Flash Realism and Texture

The authentic snapshot look of this prompt relies on four meticulously tuned engineering levers:

  1. Single Direct Smartphone Flash (lighting): Emulates the raw, unsoftened point-source illumination of a phone flash mounted adjacent to the camera sensor. The strong frontal light creates crisp micro-shadows while steep falloff plunges background roller shutters and asphalt into darkness, naturally separating the subject without synthetic background blurring.
  2. 26mm Wide-Angle Optics and Sensor Imperfections (camera): Explicitly rejecting pristine cinema lenses and creamy bokeh, the prompt specifies a 26mm equivalent smartphone focal length, handheld framing, subtle JPEG compression artifacts, and organic sensor noise in shadows, capturing the genuine social media snapshot aesthetic.
  3. Skin Texture and Flash Specular Highlights (skin): To prevent the mannequin-like surface typical of generative diffusion models, the prompt defines fine realistic texture, gentle pores, soft natural sheen. It specifically calls for specular flash reflections across the face, collarbones, and arms, grounding the image in physical lighting reality.
  4. Targeted Anti-Slop Negative Guard (negative): By prohibiting CGI, anime, studio lighting, excessive bokeh, plastic skin, over-retouched skin, the prompt preemptively suppresses the model's tendency to inject synthetic three-point lighting, glossy rendering textures, or heavy smoothing filters.

Practical Application Tips and Customization Guide

  • Execution in Chat Interfaces: When executing the prompt in ChatGPT, prepending a directive such as Generate an image strictly following this JSON specification: reinforces that the JSON block represents a rigid design schema rather than conversational dialogue.
  • Customizing Wardrobe and Locations: Creators can adapt this workflow to different fashion concepts while preserving lighting consistency by editing solely the outfit values (e.g., swapping to a leather biker jacket or oversized hoodie) and the scene field (e.g., brick alleyways or late-night convenience store storefronts).
  • Aspect Ratio and Compositional Balance: The default 4:5 vertical framing (composition) is calibrated for mobile feeds, captured from an approximate 1.8-meter working distance. This anchors an optimal mid-shot balance where the subject fills approximately 80% of the vertical frame while retaining enough dark urban environment for atmospheric context.

Original source