A Practical Guide to Prompting AI Video
Describing motion and camera work is the whole game — here's how to do it for Veo, Runway, Sora, Kling and more.
Prompting for video is not just prompting for an image with the word "video" added. A still only has to look right in one frame; a clip has to move convincingly over time. That means two ideas do most of the heavy lifting: what moves in the scene, and what the camera does. Get specific about both and your results improve dramatically.
Describe the motion, not just the scene
Image prompts describe a frozen moment. Video prompts need to describe change. Instead of "a fox in a forest," write what actually happens: "a fox trots across a mossy clearing, pauses, and turns its head toward the camera." Give the model a small piece of choreography. Vague motion ("a busy city") tends to produce either a near-static shot or chaotic, flickery movement.
Direct the camera explicitly
Camera language is the single biggest quality lever in AI video. Tools understand terms borrowed straight from filmmaking:
- Pan / tilt — the camera rotates left/right or up/down from a fixed spot.
- Dolly / push in / pull out — the camera physically moves toward or away from the subject.
- Tracking / follow shot — the camera moves alongside a moving subject.
- Crane / aerial — sweeping, elevated movement.
- Static / locked-off — no camera movement, which is often the safest choice for a clean result.
"Slow dolly-in on the fox as it turns toward camera" tells the tool exactly how the shot should feel.
One clear action per clip
Most current models produce a few seconds at a time. Trying to cram a whole sequence — "she walks in, sits down, opens a book, and it starts raining" — usually produces mush. Give each clip one clear action and one camera move, then stitch clips together afterward if you need a sequence.
Match the phrasing to the tool
Different tools reward slightly different prompt styles:
- Veo, Runway, Sora respond well to natural, cinematic prose — describe the shot like a sentence from a screenplay, including camera direction.
- Kling likes a descriptive prompt plus a distinct camera-movement cue, and supports a separate negative prompt.
- Wan works best with a structured, six-part prompt: camera movement, subject/scene, motion, camera language, style, and atmosphere.
- LTX prefers chronological, literal descriptions — main action first, then motion detail, then environment and camera last.
Set the atmosphere and pace
Tell the model the mood and the tempo. "Calm, dreamy, slow motion" and "energetic, fast-paced, handheld" pull a clip in completely different directions even with the same subject. Lighting and time of day matter here just as much as they do for stills.
A workable template
A dependable pattern is: [camera movement] + [subject] + [specific action] + [setting] + [lighting/mood]. For example: "Slow tracking shot following a cyclist along a coastal road at sunset, warm golden light, calm and cinematic." It's one action, one camera move, and a clear mood — exactly what these models handle best.
Our Prompt Generator has a video mode that structures prompts the way each of these tools prefers, so you don't have to remember every convention.