Last updated: 2026-07-13
Text-to-video and image-to-video models in 2026 (Veo 3.1, Kling 3.0, Runway Gen-4.5, Sora 2) are good enough to ship, but only if you direct them like a cinematographer, not a search engine. The craft is no longer "describe a pretty scene" — it's writing a shot with one clear camera move, locking your subject across clips, and assembling short generations into something that reads as continuous. This chapter gives you a per-shot prompt formula, the discipline rules that separate usable clips from slot-machine rerolls, and a realistic assembly pipeline.
What matters most
- Every strong clip prompt fills five slots in this order: Cinematography (shot type + camera move) → Subject (specific description) → Action (one continuous action) → Context (location, time, lighting) → Style & Ambiance (film stock, mood, grade). This is Google's own recommended Veo structure and it transfers to every model.
- Front-load the camera. Models weight the start of the prompt heavily, so lead with the shot framing and move ('Slow dolly-in, medium shot of...') rather than burying it at the end.
- Be concrete on the subject: 'a woman in her twenties with wavy brown hair and light freckles, red wool coat' beats 'a stylish woman.' Vague identity is the #1 cause of drift between shots.
- Name lighting explicitly — 'harsh fluorescent overhead + green monitor glow,' 'warm low-key single key from frame left.' Lighting direction and color temperature secretly anchor how the model reconstructs faces and materials.
- Add audio deliberately on models that generate it (Veo, Kling, Sora 2): put spoken lines in quotes, tag SFX ('SFX: distant thunder'), and set ambience ('Ambient: quiet hum of a starship bridge'). Undefined audio gets filled with generic noise.
- Shot type vocabulary that parses reliably: wide/establishing, medium, two-shot, close-up, extreme close-up, low angle, high angle, over-the-shoulder, POV, reverse shot.
Common mistakes to avoid
- Don't stack five adjectives of 'cinematic/epic/beautiful' — they add nothing and crowd out the concrete direction the model can act on.
- Don't describe two lighting setups or two locations in one clip; the model will blend them into mush ('concept bleed').
- Avoid negatives ('no blur, not cartoonish') on most models — they often summon the thing you're excluding. Describe what you DO want instead.
- Two simultaneous moves ('zoom while panning while craning') is the most common cause of the 'liquid/rubber' look — pick one.
The short version
- Write shots, not scenes: Cinematography → Subject → Action → Context → Style, with the camera move front-loaded and lighting/audio named explicitly.
- One clip = one camera move + one subject action. Anything more splits into multiple clips and gets cut together in post.
- Consistency comes from features + frames, not prose: use each model's reference/Ingredients/cameo feature, chain a sharp exported frame into the next shot, and lock the identity wording verbatim.
- Match the dialect to the model: Veo = structured/reference-anchored, Kling = time-coded script, Runway = physics/force verbs + camera tokens, Sora 2 = causal 'why it happens' language.
- Model snapshot (Jul 2026): Veo 3.1 for hero/atmosphere + native audio; Kling 3.0 for value and long clips; Runway Gen-4.5 for edit tools and iteration; Sora 2 for physics and likeness cameos — verify specs before quoting, they change monthly.
This is one lane of the full system. Get all ten — prompt skeletons, copy-paste templates, worked examples, and the 2026 tool picks — in The AI Creator's Playbook: get the complete 70-page playbook ▸