The text-to-video prompting guide that actually works
Camera language, lighting cues, and pacing — the practical structure behind prompts that produce cinematic, controllable shots.
A great AI video prompt reads like a shot list, not a wish. The models respond to the vocabulary of cinema — so the more you speak in shots, lenses, and light, the more control you get.
Start with the shot, not the story
Open every prompt by naming the shot type and camera move. "Slow dolly-in", "handheld tracking shot", "static wide" — these set the entire frame before you describe a single object. The subject comes second.
- Shot size: extreme close-up, close-up, medium, wide, establishing
- Camera move: dolly, pan, tilt, crane, orbit, handheld
- Lens feel: 24mm wide, 85mm portrait, shallow depth of field
- Pace: slow, deliberate, snappy
Light is half the prompt
Cinematic output lives and dies on lighting language. "Golden hour", "hard rim light", "soft overcast", "neon practicals on wet asphalt" — each pushes the model toward a specific mood and contrast curve. Vague prompts get vague light.
“Describe the light first and the subject will follow.”
Control motion with verbs
Models infer motion from your verbs. "Drifting", "racing", "billowing", "settling" each carry a speed and weight. Pair the subject verb with the camera verb so the two motions complement rather than fight each other.
Iterate one variable at a time
When a clip is close but not right, change one thing — the lens, the light, or the pace — and regenerate. Changing everything at once makes it impossible to learn what the model responded to. Treat it like grading a shot, not rolling dice.
Ready to make something?
Turn your next idea into a cinematic clip in the Omnira studio.
Open the Studio