Mastering AI Video Prompts: Sora, Runway Gen-3 & Kling AI Motion Guide
Camera motion keywords, lighting cues, temporal consistency tricks, and exact prompt templates to produce cinematic AI videos.
Introduction
Text-to-video models in 2026 are no longer judged on whether they can produce a plausible clip. They are judged on whether a creator can reproduce a specific shot on demand. The gap between a lucky generation and a repeatable one comes down to prompt structure. A vague line such as "a woman walks through a city at night" gives the model a thousand valid interpretations, and each regeneration picks a different one. A structured prompt that fixes subject, lens, camera path, lighting and temporal behavior narrows that space until the output becomes predictable.
That is the practical meaning of searches for sora prompts, runway gen 3 prompts and kling ai video prompts. People are not looking for magic phrases. They want a formula they can apply to any scene and a vocabulary the model understands. This guide gives you both: how video diffusion models read a prompt, a comparison of how Sora, Runway Gen-3 and Kling respond to the same instructions, and a reusable template with working code for batch generation.
One recurring theme in recent industry discussion is cost and reliability. A Hacker News thread titled "Claude Status – Elevated errors for multiple models" drew over 100 upvotes, a reminder that any generation pipeline needs retry logic and fallbacks. Video is far more expensive per call than text, so the prompt discipline described below pays for itself in fewer wasted renders.
Architectural Breakdown & Core Mechanics
Modern video generators are latent diffusion transformers. They compress video into a spatiotemporal latent space, then denoise patches that span both space and time. Your prompt is encoded by a text encoder into conditioning vectors that steer every denoising step. Three consequences follow for prompt writers.
1. Order and density matter. Text encoders weight early tokens more strongly. Put the subject and the primary action first, then camera, then environment, then style. Overloading a prompt with 15 competing adjectives dilutes each one.
2. Motion is described, not implied. Image models only need a static scene. Video models need an explicit statement of what moves, how fast, and in which direction. If you do not specify camera motion, the model chooses one, usually a slow drift, and it may change between regenerations.
3. Temporal consistency is a conditioning problem. Characters drift because the model re-derives their appearance every few frames from a loosely worded description. Fixed, concrete attributes (hair, clothing color, a single distinguishing item) give the attention layers stable anchors.
The five-layer prompt formula
Every strong video prompt in this guide uses the same five layers:
- Subject and action: who or what, doing exactly one primary thing.
- Camera: shot size, lens feel, and a single movement verb.
- Environment and time: location, weather, time of day.
- Lighting and color: source, quality, palette.
- Style and format: film stock feel, aspect ratio, duration cues.
Camera movement vocabulary
These keywords are widely understood across all three models. Use one movement per shot, because stacking several tends to produce jittery or contradictory motion.
| Keyword | Effect | Best used for |
|---|---|---|
| Static locked-off shot | No camera motion, subject moves | Dialogue, product reveals |
| Slow dolly in | Camera pushes toward subject | Emotional emphasis |
| Dolly out / pull back | Camera retreats, widens context | Reveals, endings |
| Tracking shot | Camera follows subject laterally | Walking, vehicles |
| Orbit / arc shot | Camera circles subject | Hero product shots |
| Crane up / jib up | Vertical rise | Establishing shots |
| Handheld | Subtle organic shake | Documentary realism |
| FPV drone dive | Fast flying perspective | Landscapes, action |
| Rack focus | Focus shifts between planes | Narrative emphasis |
Lighting cues that survive the encoder
Concrete, physical lighting terms work better than mood words. "Low-angle golden hour sunlight with long shadows" outperforms "beautiful warm light". Other reliable cues: "overcast soft diffused light", "neon signage reflected in wet asphalt", "single practical lamp, deep shadows", "backlit with rim light and haze". If you already know how lighting phrasing steers still images, the same ideas carry over. The Midjourney v4/v5 Cinematic Volumetric Lighting Portrait prompt is a good reference for volumetric vocabulary you can adapt to motion.
Comparative Benchmarks & Evaluation Matrix
The table below compares qualitative behavior observed when feeding the same structured prompts into each tool. These are relative ratings from hands-on testing patterns, not published benchmark figures, and model versions change often, so re-test before committing to a pipeline.
| Capability | Sora | Runway Gen-3 | Kling AI |
|---|---|---|---|
| Prompt adherence on long descriptions | Strong | Moderate to strong | Strong |
| Camera movement obedience | Strong | Strong, favors explicit verbs | Strong |
| Physical plausibility (cloth, liquids, collisions) | Strong | Moderate | Strong |
| Character consistency across clips | Moderate | Moderate, improves with reference images | Moderate to strong |
| Stylized and animated looks | Strong | Strong | Moderate |
| Sensitivity to prompt length | Tolerates long prompts | Prefers concise, ordered prompts | Tolerates long prompts |
| Best prompt style | Narrative paragraph | Comma-separated shot spec | Descriptive with explicit motion |
| Iteration speed | Slower | Fast | Moderate |
| Image-to-video control | Good | Very good | Very good |
How to read the matrix
Sora rewards a narrative paragraph that reads like a shot description from a screenplay. Runway Gen-3 responds best to a tight, ordered shot specification, so trim adjectives and keep camera verbs explicit. Kling handles detailed physical action well, which makes it a strong choice for scenes with water, fabric or crowds.
For consistent characters, no model is reliable from text alone. The dependable workaround is the same in all three: generate a hero still first, then use image-to-video so the first frame fixes the character, and keep the text prompt focused on motion only.
Practical Step-by-Step Implementation Recipe
This recipe turns the five-layer formula into a reusable generator, so every shot in a project shares the same style block and character anchors.
Step 1: Define a character and style lock
Write the character description once, using concrete and visually checkable attributes. Reuse it verbatim in every prompt.
Step 2: Build prompts from structured data
from dataclasses import dataclass
@dataclass
class Shot:
subject_action: str
camera: str
environment: str
lighting: str
style: str = "35mm film look, shallow depth of field, 16:9"
CHARACTER = (
"a woman in her thirties with short silver hair, "
"wearing a mustard yellow raincoat and black boots"
)
STYLE_LOCK = "muted teal and amber color grade, fine film grain"
def build_prompt(shot: Shot, target: str = "sora") -> str:
parts = [
shot.subject_action.format(character=CHARACTER),
shot.camera,
shot.environment,
shot.lighting,
STYLE_LOCK,
shot.style,
]
if target == "runway":
# Runway Gen-3 prefers a compact, comma-separated spec
return ", ".join(p.rstrip(".") for p in parts)
# Sora and Kling accept fuller sentences
return ". ".join(p.rstrip(".") for p in parts) + "."
shots = [
Shot(
subject_action="{character} walks slowly across a wet crosswalk",
camera="Tracking shot at eye level, 35mm lens feel, steady motion",
environment="Empty downtown street at night after rain",
lighting="Neon signage reflected in puddles, soft rim light from behind",
),
Shot(
subject_action="{character} stops and looks up at a flickering sign",
camera="Slow dolly in from medium shot to close-up",
environment="Same street, light drizzle falling",
lighting="Intermittent pink and cyan neon flicker on the face",
),
]
for i, shot in enumerate(shots, start=1):
print(f"--- Shot {i} (Runway) ---")
print(build_prompt(shot, target="runway"))
print(f"--- Shot {i} (Sora) ---")
print(build_prompt(shot, target="sora"))Step 3: Submit with retries and a seed policy
Video generation is expensive and APIs fail. Wrap every submission in exponential backoff and log the exact prompt and parameters so you can reproduce a good result.
import json
import time
import random
import requests
def generate_with_retry(endpoint: str, api_key: str, payload: dict,
max_attempts: int = 4) -> dict:
headers = {"Authorization": f"Bearer {api_key}"}
for attempt in range(1, max_attempts + 1):
try:
resp = requests.post(endpoint, json=payload, headers=headers, timeout=60)
if resp.status_code == 429 or resp.status_code >= 500:
raise RuntimeError(f"transient error {resp.status_code}")
resp.raise_for_status()
result = resp.json()
with open("generation_log.jsonl", "a") as log:
log.write(json.dumps({"payload": payload, "result": result}) + "\n")
return result
except (requests.RequestException, RuntimeError) as exc:
if attempt == max_attempts:
raise
delay = (2 ** attempt) + random.uniform(0, 1)
print(f"Attempt {attempt} failed ({exc}); retrying in {delay:.1f}s")
time.sleep(delay)
raise RuntimeError("unreachable")Endpoint paths and payload field names differ per provider, so check each vendor's current API documentation before wiring this in. The logging and retry pattern is the part that transfers.
Step 4: Fix character drift with a hero frame
- Generate a still of the character with an image model.
- Pick the best frame and use it as the first-frame input for image-to-video.
- Keep the text prompt to motion and camera only: "She turns her head toward the camera, slow dolly in, neon flicker."
- For each following shot, reuse the same hero frame or a clean frame extracted from the end of the previous clip.
Step 5: Iterate one variable at a time
Change a single layer per regeneration. If the camera move is wrong, edit only the camera line. Changing three layers at once makes it impossible to learn what worked.
Ready-to-use templates
Cinematic establishing shot:
Crane up over a coastal town at dawn, wide aerial view, mist rolling between rooftops, soft pink sunrise light with long shadows, anamorphic lens flare, calm slow motion, 16:9
Product orbit:
Slow 360-degree orbit around a matte black wireless speaker on a wet slate surface, single softbox key light with a blue rim light, shallow depth of field, macro lens feel, clean studio background
Handheld documentary:
Handheld shot following a street vendor assembling a stall at sunrise, subtle camera shake, natural overcast light, steam rising from a cart, documentary realism
Negative guidance (where supported):
morphing faces, extra fingers, flickering textures, text artifacts, sudden scene cuts
Strategic Catalog Integrations
Video prompting does not live in isolation. A practical workflow chains several tools, and the AIFuller catalog covers most of the links.
Concepting and shot lists. Use ChatGPT or Claude to expand a one-line idea into a numbered shot list, with each shot written in the five-layer format. Gemini is useful when you want to feed it a reference clip or storyboard images and ask for a prompt that reproduces the look.
Hero frames for character lock. Generate the first frame with Midjourney for premium photorealism, or with Flux.1 if you want an open model you can run and fine-tune yourself. If you prefer local control, the AUTOMATIC1111 Stable Diffusion WebUI lets you iterate on stills without per-image costs.
Style references from the prompt library. Aesthetic shorthand transfers well to video. The 1980s Neon Synthwave Grid & Sports Car and 1980s Japanese City Pop Anime Rooftop prompts give you tested palettes and composition cues, and the Wes Anderson Symmetrical Pastel Boutique Hotel prompt is a strong template for locked-off symmetrical shots. The Studio Ghibli Grand Bathhouse & Spirit Train (Spirited Away) prompt works as a base for animated looks. Browse community libraries for community prompt patterns you can adapt.
Running generation locally. A Hacker News thread, "Frontier AI on Your Own Hardware", drew more than 150 upvotes, which reflects growing interest in local pipelines. Ollama handles the text side, such as a local model that drafts and rewrites shot prompts offline, while image and video models run through their own runtimes.
Automating the pipeline. If you are building a batch tool around the Python code above, the Claude Code Senior Staff Engineer Protocol rule set keeps an agent making atomic commits and verifying its work, and the FastAPI and Pydantic rules help if you expose the generator as a service with a job queue. To go deeper on orchestration, the Full Stack LLM Bootcamp & Production Agents and Building Systems with ChatGPT API courses cover the structured prompting and retry patterns used here.
Frequently Asked Questions (FAQ)
What is the best structure for Sora prompts?
Write a short narrative paragraph in the order subject and action, camera, environment, lighting, then style. Keep one primary action and one camera movement per shot. Sora tolerates longer descriptions than most models, but extra detail only helps when it is concrete and visual rather than emotional.
How do Runway Gen-3 prompts differ from Kling AI video prompts?
Runway Gen-3 does best with a compact, ordered shot specification where camera verbs are explicit and adjectives are trimmed. Kling AI handles more descriptive text and complex physical action well, so you can spell out fabric, water or crowd behavior in detail. In both, image-to-video gives you more control than text alone.
How do I keep a character consistent across multiple clips?
Text-only consistency is unreliable in current models. Generate a single hero image, reuse it as the first frame for each clip, repeat the same concrete character description in every prompt, and keep style words identical across shots. Extracting the final frame of one clip as the first frame of the next also reduces drift.
Which camera movement keywords work most reliably?
Dolly in, dolly out, tracking shot, orbit, crane up and static locked-off shot are the most consistently obeyed across tools. Combine only one movement per shot. If a clip looks jittery, remove secondary motion words before changing anything else.