AI video prompting: shots that actually cut together

Camera, motion, and consistency rules for Veo, Kling, Runway and Sora — and the one-move rule that kills the jitter.

Last updated: 2026-07-13

Text-to-video and image-to-video models in 2026 (Veo 3.1, Kling 3.0, Runway Gen-4.5, Sora 2) are good enough to ship, but only if you direct them like a cinematographer, not a search engine. The craft is no longer "describe a pretty scene" — it's writing a shot with one clear camera move, locking your subject across clips, and assembling short generations into something that reads as continuous. This chapter gives you a per-shot prompt formula, the discipline rules that separate usable clips from slot-machine rerolls, and a realistic assembly pipeline.

What matters most

  • Every strong clip prompt fills five slots in this order: Cinematography (shot type + camera move) → Subject (specific description) → Action (one continuous action) → Context (location, time, lighting) → Style & Ambiance (film stock, mood, grade). This is Google's own recommended Veo structure and it transfers to every model.
  • Front-load the camera. Models weight the start of the prompt heavily, so lead with the shot framing and move ('Slow dolly-in, medium shot of...') rather than burying it at the end.
  • Be concrete on the subject: 'a woman in her twenties with wavy brown hair and light freckles, red wool coat' beats 'a stylish woman.' Vague identity is the #1 cause of drift between shots.
  • Name lighting explicitly — 'harsh fluorescent overhead + green monitor glow,' 'warm low-key single key from frame left.' Lighting direction and color temperature secretly anchor how the model reconstructs faces and materials.
  • Add audio deliberately on models that generate it (Veo, Kling, Sora 2): put spoken lines in quotes, tag SFX ('SFX: distant thunder'), and set ambience ('Ambient: quiet hum of a starship bridge'). Undefined audio gets filled with generic noise.
  • Shot type vocabulary that parses reliably: wide/establishing, medium, two-shot, close-up, extreme close-up, low angle, high angle, over-the-shoulder, POV, reverse shot.

Common mistakes to avoid

The short version


This is one lane of the full system. Get all ten — prompt skeletons, copy-paste templates, worked examples, and the 2026 tool picks — in The AI Creator's Playbook: get the complete 70-page playbook ▸

Found this useful? Get a new guide every week

One email a week: a new guide, a new track, one good prompt. No spam — unsubscribe is one reply.

🎓 Practice this interactively: this guide has a self-graded course in Learn Lab.Start practicing →
← How to write AI image prompts that actually woBuild a self-improving AI system (the method b →