The structure that works
Image models pay the most attention to the first few words of a prompt. Lead with what matters most. A good order: Subject → Composition → Lighting → Environment → Mood → Style “A weary detective in a rumpled trench coat leans against a rain-slicked lamppost under a single pool of amber light, noir atmosphere, shot on 35mm film” beats “detective, trench coat, rain, lamppost, noir, cinematic.” Write in flowing natural language — not a tag list.Be specific about the subject
“A woman in a coat” gives the model nothing to work with. “A tall woman in her thirties with short auburn hair and angular features, wearing a rumpled navy peacoat” gives it a character. For Actors especially, specify:- Age range (early thirties, mid fifties)
- Ethnicity or features (East Asian, high cheekbones, freckled)
- Build (tall and lean, stocky, petite)
- Hair (colour, length, style)
- Clothing (material, cut, colour — “rumpled wool overcoat” not “coat”)
- Expression and pose (slight smirk, arms crossed, shoulders forward)
Composition and camera
Use filmmaking vocabulary — the models are trained on it. Shot size: extreme close-up, close-up, medium shot, wide shot, establishing shot. Be explicit about framing. Angle: eye level, low angle, high angle, Dutch angle, over-the-shoulder. Depth of field: “Shallow depth of field, subject sharp, background bokeh” or “deep depth of field, everything in focus.” Lens references: “Shot on Canon EOS R5 with an 85mm lens,” “wide-angle 24mm,” “anamorphic lens.” Lens choice bundles sharpness, compression, and depth-of-field behaviour into one phrase.Lighting
Lighting shapes mood more than almost anything else. Always specify it — unspecified lighting defaults to flat. Direction: key light from camera-left, backlit, side-lit, top-lit. Named patterns: Rembrandt lighting, butterfly lighting, split lighting. These produce distinct, recognisable results. Natural light: golden hour, blue hour, overcast, harsh midday sun, cold moonlight. Practicals: neon glow, single bare bulb, candlelight, screen glow, fireplace light. Atmosphere: film grain, haze, fog, dust motes in light beams, rain-slicked surfaces. These add depth.Style
A few words of style direction pull the whole image toward the look you want. Embed them in the prompt naturally.Photoreal
“Shot on Canon EOS R5,” “Fujifilm XT4 with 56mm f/1.2,” “Hasselblad medium format.” Bundles colour science, sharpness, and depth-of-field behaviour.
Cinematic
“Cinematic” is the reliable workhorse. Combine with “film grain,” “anamorphic lens,” “rich colour grading,” “dramatic lighting.”
Painterly
“Oil painting,” “watercolour illustration,” “charcoal sketch,” “impressionist glow,” “lush brushstrokes.”
Stylised
“Studio Ghibli style,” “cel-shaded animation,” “1970s sci-fi book cover art,” “Pixar 3D render.” Each phrase bundles palette, composition, and texture.
Aspect ratio
Pick the ratio before you generate — reshaping afterwards loses quality.- 16:9 — widescreen, most cinematic scenes
- 4:3 — classic film, TV, documentary feel
- 9:16 — vertical, mobile-first, hero portraits
- 1:1 — square, social media
What to avoid
- Quality keyword stacking. “Masterpiece, best quality, 8k, ultra-detailed, award-winning” does nothing. Modern models default to high quality. Every word should carry visual information.
- Meta-language. “A beautiful image of,” “an amazing photograph of,” “stunning.” Zero visual content. Cut it.
- Contradictions. “Bright well-lit room with deep shadows” confuses the model. Pick one.
- Text in images. Most image models render text as gibberish. If you need a sign or label, generate the image without it and composite text later.
- Tight close-ups on hands. Hands are the classic AI failure mode. Frame wider, or pose hands in pockets, behind the back, or holding a simple object.
- Over-complex compositions. Four-plus characters in one frame drift fast. Keep it to two or three, and use spatial blocking: “A on the left facing right; B on the right facing left.”
Keeping characters consistent
If you reshot an Actor in scene three and it doesn’t match the one from scene one, the problem is usually wording drift. Lock the description. “Grey wool overcoat” must not become “dark coat” four frames later. Ozu stores your Actor’s hero look and wardrobe variants and feeds them as references into downstream frames. Keep hero looks clean and well-lit — they set the baseline for every frame that uses that Actor.Related
Video Prompts
Writing motion, camera moves, and pacing on top of a still frame.
Actors and Wardrobe
Hero Looks, Wardrobe Variants, and how identity carries across scenes.
Sets
Environment and location assets — consistency across shots.
Safety Blocks
Why a generation might be blocked and how to reshot around it.