← Community library
Realism guide
How to make generated video look filmed, not generated: camera, light, color, people, props, product fidelity and the common failure modes.
Guide by Takeloom
Sign up to use it# Production Realism Prompt Guide
This document is injected context for the LLM when writing video generation prompts for the video model (for example Seedance 2.0).
Follow every rule here when constructing a prompt. These are distilled from iterative testing and represent the most reliable way to produce professional, non-AI-looking commercial video output.
**Scope: this guide governs visual rendering only** — camera work, light, color, scene composition, production quality. It says nothing about pacing, audio, or energy: a script can (and for ads, should) have cuts, continuous voiceover, and a music bed while every frame still obeys these rules. Never use "production quality" as a reason to make a video quiet or slow — that is governed by the TV ad guide.
**The target look:** a professionally shot 16:9 television commercial — real cinema glass, motivated lighting, a deliberate grade — that still reads as *filmed*, not *generated*. The video model's default output overshoots into AI-gloss: hyper-saturated HDR color, plastic skin, rim-light halos, impossibly perfect surfaces, floaty physics. Your job is to land on broadcast-real, which sits *below* the model's default level of polish.
**Important: this guide teaches techniques, not a scene.** Every example below is an illustration of the *level of physical specificity* required — never a template to copy. The scene itself (location, time of day, objects, people) must be derived fresh from the brand and product at hand (see "Scene Selection" below). Do not reuse a scene from an example just because it appears in this guide.
**Not only for ads.** This guide was written for TV commercials and most of its examples are products. Every rule about camera, light, color, people, scene composition, invented text and anatomy applies to every kind of video — short films, music videos, explainers, reels. The rules about the product (label, fidelity, mechanics, usage logic) apply whenever a product or a key prop is on screen; in a video without a product, apply the same care to the props the story depends on.
---
## Scene Selection — Derive the Scene from the Product
Before writing any prompt, choose the scene from the product's natural usage context — the place and moment a real customer would actually use or encounter it. Ask:
1. **Where is this product genuinely used?** A condiment lives at meals; a shampoo in a bathroom; a camera at a birthday, a trailhead, a football pitch; a detergent in a laundry room.
2. **What recognizable human moment involves it?** TV ads are staged, but the staging depicts a moment the viewer instantly recognizes from their own life: cooking dinner, the kids' bath-time chaos, packing the car at dawn. Specificity of the moment is what makes a commercial feel true.
3. **What would this look like shot for broadcast?** A director would pick the angle that tells the story, light it to look like the real place at that hour, and dress the set to feel lived-in — not sterile, not showroom.
Vary scenes across generations for the same brand — indoor/outdoor, day/evening, kitchen/car/park/pitch. If two consecutive prompts landed in the same setting, deliberately pick a different one. Never fall back to a "default" location.
---
## How the Video Model Reads Prompts — The Foundational Rule
**Describe causes, not outcomes.**
The model interprets outcome language using its default aesthetic — hyper-produced, over-graded, AI-looking. When you describe the physical causes (light source position, lens and camera support, exposure logic, grade discipline), the model is forced to render the physics you specified, and the aesthetic result follows naturally.
| Instead of (outcome) | Write this (cause) |
|---|---|
| "Warm afternoon light" | "Sun at 25° above horizon from camera-left — warm cast on horizontal surfaces only" |
| "Cinematic shallow depth of field" | "Digital cinema camera, 50mm lens at T2.8, 1.2m from subject — background softly out of focus, subject fully sharp" |
| "Epic cinematic feel" | Never write this. Specify camera support, lens, light source, and grade discipline instead. |
| "Rich saturated colors" | Never write this. Describe the light source and the restrained grade instead. |
| "Beautiful dreamy slow motion" | Never write this. If you need emphasis, write the physical action at natural speed. |
---
## Prompt Structure
### Opening line — always start with the capture context
The very first line sets the model's entire default. Use it to establish "professionally filmed television commercial" before anything else — the production context, not an aesthetic adjective:
```
A [format] television commercial (or short film, music video, explainer), shot on a digital cinema camera with real lenses and practical, motivated lighting — filmed footage, not CGI.
```
Never open with aesthetic language ("A breathtaking, cinematic shot of..."). Never describe the video as "ultra-realistic" or "8K" — the model's version of those words is the AI-gloss look this guide exists to prevent.
### Reference image declarations
Always declare what each reference image is and what role it plays — immediately after the opening line:
```
@reference_image[0] is the [product] — replicate its label, shape, and colors exactly.
@reference_image[1]: use ONLY for [environment layout / spatial depth / architectural arrangement]. Do NOT copy any people, food, drinks, objects, or color treatment from this image.
@reference_image[2]: use ONLY for [surface setup style / room layout]. Do NOT copy any objects or items visible in this image.
```
### Time segments
Use explicit time blocks. Describe what is physically in frame and what is physically happening — concrete actions and objects, not mood words:
```
[0s-3s]: [What is in the frame. What happens physically.]
[3s-6s]: [What changes. What the product does. Where gaze goes.]
```
### Character limit
Every video model has a **prompt character limit** (given to the agent as the prompt limit). The pipeline reserves headroom for the text-overlay block — check length before submitting.
---
## Camera — Most Important Variable
The target is disciplined professional camera work. Two failure directions exist: the model's default *floaty, over-animated* camera (constant drifting, orbiting, speed-ramping — the #1 AI tell), and accidentally prompting *amateur* footage. Override both explicitly.
### Always specify
```
CAMERA: Digital cinema camera, [lens focal length] lens, [distance] from subject, [angle], [height].
[Support: locked on tripod / smooth dolly / gimbal — pick one per scene.]
[Either: "Camera locked, no movement." or ONE motivated move: "Slow 10cm push toward the subject over the full scene, constant speed." ]
Exposure set for the subject; highlights roll off naturally, nothing clipped, nothing lifted.
```
### Key rules
- **One camera behavior per scene.** Locked-off, or a single slow motivated move (push, pan with the action, short dolly). Never stack moves, never orbit, never speed-ramp. Unmotivated constant drift is the strongest AI tell in generated video.
- **Composed framing.** State the composition intent: rule-of-thirds placement, headroom, leading room in the direction of action. TV frames are deliberate — but state it physically ("subject on the left third, looking right into open frame"), not as "beautifully composed."
- **Real lens behavior.** Give focal length and distance so depth of field has a physical cause. Focus stays on the subject; no auto-refocus hunting, no rack focus unless a scene explicitly needs it (at most once per spot).
- **Natural motion speed.** All action at natural speed. No slow motion unless the script explicitly calls for it in one scene, and then name the frame rate cause ("shot at 120fps, played back at 24fps").
---
## Lighting — Physics-First Description
### Core rule
Professional but motivated: every light in the scene must plausibly come from something that exists in that world. Describe: source position → what it hits → what it does NOT hit.
```
LIGHT: Key source: [the scene's plausible dominant source — window daylight, sun, practical lamps] from [direction], [height/angle].
Soft fill from [bounce/ambient], 2–3 stops below key — shadow side retains detail but stays visibly darker.
Shadows fall to [direction], soft-edged, anchored to their objects.
No glow, no halo, no light wrapping around subjects from nowhere.
```
### The halo rule
The model's default "looks good" lighting adds a glowing rim/halo around subjects with no source behind them — a major AI tell. A *motivated* backlight (a window behind the subject, low sun) is fine television craft; an unmotivated glow is not. Always state where any backlight comes from, and otherwise prohibit it:
```
No rim light or edge glow unless it comes from the named source behind the subject.
```
### Example formulas (adapt the source to the scene — do not default to one)
**Outdoor, late afternoon:**
```
Key: sun at 20–25° above horizon from camera-left, warm but not orange.
Shadow side lit only by sky bounce — visibly darker, detail retained.
Sky in frame reads bright but not blown out.
```
**Indoor, daytime:**
```
Key: daylight through a window camera-right — soft, directional falloff right to left.
Interior practicals off. Left side of the scene 2 stops down, natural bounce only.
Window area bright but with visible exterior detail, not clipped white.
```
**Indoor, evening:**
```
Key: warm practical lamps in frame (table lamp, kitchen pendant) doing the actual lighting.
Pools of light around each practical, falloff to dim between them. Skin stays natural, not orange.
```
### What not to write
- "Golden hour glow" → triggers orange saturation and halo lighting
- "Dramatic lighting" / "moody lighting" → model invents unmotivated sources
- "Soft, beautiful light" → outcome language; name the source instead
- "Studio lighting" → produces the sterile packshot look unless the scene IS a studio
---
## Color Rendering
The model's default is hyper-graded — punchy, saturated, HDR-tone-mapped, teal-orange. A broadcast commercial is graded, but with restraint. Override the default explicitly:
```
COLOR: Clean broadcast grade — accurate, natural skin tones; true-to-life colors with restrained saturation.
No HDR tone mapping, no teal-orange split, no crushed blacks, no lifted matte shadows.
The product's label colors match the reference image exactly.
Identical color treatment for foreground and background — no differential processing.
```
Skin tone accuracy is the anchor: if skin reads natural, the whole grade reads real. The foreground/background differential matters too — the model often applies heavier processing to foreground elements than to background. Explicitly requiring consistent treatment across the entire frame prevents this.
---
## Reference Images — Three Failure Modes to Prevent
### 1. Color grade bleed
The model treats reference images as full color references by default. A stylized vibe image will drag its color grade into the entire scene.
Fix: explicitly restrict each reference image's role and ban its color treatment:
```
@reference_image[1]: use ONLY for [environment] layout and spatial depth.
Do NOT use its color saturation, contrast, or color grade — ignore the color treatment entirely.
```
### 2. Object copying
The model copies specific objects from reference images into the scene. "Match the energy" is not enough — if the reference image has soda bottles and paper cups, those will appear.
Fix: ban copying, then describe only what you want:
```
Do NOT copy any food, drinks, bottles, cups, bowls, or objects from @reference_image[1] or @reference_image[2].
```
Then list the exact items you want and end with "Nothing else."
**This also means: reference images can drag in a wrong *scene*.** If a vibe image shows a rooftop and the chosen scene is a kitchen, either use a reference image that matches the chosen scene or drop the vibe reference entirely and describe the environment in text. Never let a leftover reference image from a previous brand decide the location.
### 3. Foreground saturation bleed
Even after restricting color grade, reference image colors can bleed into the foreground. The explicit color rendering block (see above) addresses this.
### Selecting reference images
Choose vibe/environment reference images based on **spatial composition and layout matching the chosen scene**, not color realism. A heavily graded editorial photo with the right layout is a better reference image than a flat photo with the wrong composition — provided you suppress its color grade in the prompt. The suppression language is the control mechanism; the composition is what the image actually contributes.
### Privacy filter (some providers, e.g. ByteDance Seedance)
Reference images containing close-up human faces are blocked with `InputImageSensitiveContentDetected.PrivacyInformation`. Distant background figures are generally fine. Crop reference images to remove close-up faces before use. Never pass a real person's photo as a reference image — describe them in text only.
---
## Scene Composition — Explicit Item Listing
Directional language ("lived-in kitchen," "cozy," "authentic feel") is overridden by reference image defaults. Only explicitly named items reliably appear.
```
# Doesn't work:
"A lived-in family kitchen"
# Works:
"On the kitchen counter: a wooden cutting board with half-chopped vegetables,
a kid's drawing taped to the fridge door at a slight angle, two mismatched mugs
near the sink, a dish towel draped over the oven handle. Nothing else."
```
**"Nothing else" matters** — it signals the list is exhaustive, not partial.
### Set dressing that reads as real
Professional art departments make sets look *inhabited*, not sterile. The model defaults to showroom-perfect. Invent the equivalent for the chosen scene — examples of the category, not a fixed list:
- Objects at believable angles, not aligned to the camera
- Traces of the activity in progress: an open lens cap, a jacket over a chair back, a half-finished drink
- Texture and age where the world would have it: a scuffed doorframe, a worn tabletop
- The product placed with intent (it IS the hero) but inside a world that doesn't revolve around it geometrically
---
## Subject / Person
### The plastic-skin problem
The model defaults to airbrushed, poreless, symmetrical AI faces. TV commercials cast attractive, relatable people — but they are *real* people with skin texture. Override explicitly:
```
PERSON: [Age, build, ethnicity if relevant]. Natural skin with visible texture and pores —
no airbrushing, no wax-smooth skin. [Specific everyday wardrobe appropriate to the scene.]
[Specific action involving the product or scene.] Expression arises from the action, not posed at camera.
```
Key rules:
- **Give every person an action.** A person doing something (framing a shot, lifting a child, pouring) renders far more believably than a person existing to be looked at.
- Gaze at camera only if the script is deliberately direct-address; otherwise gaze follows the action.
- Casting-brief adjectives ("confident," "effortless," "radiant") produce stock-ad mannequins — describe wardrobe, age, and action instead.
- If a scene works with hands/body only (product close-ups, demos), prefer it — it removes the face-rendering risk entirely and reads as premium product photography.
---
## Product Visibility
### Label facing camera
The label drifts away from camera unless reinforced in every time segment:
```
[0s-3s]: @reference_image[0] on the [surface], label facing toward camera.
[3s-6s]: [Person] lifts @reference_image[0] — label visible as it tilts toward camera.
```
### Product action visibility — be physically explicit
"They use the product" does not reliably produce a visible action. Describe the physical path of whatever the product does — pour, spray, click, focus, application:
```
# Pour example:
Thick [color] [substance] streams out of the nozzle — a visible, continuous stream
in the air between the container and [the target], landing on [exact spot].
The stream is [color] and opaque. Natural speed, not slow-motion.
# Camera-product example:
Her thumb half-presses the shutter; the focus box snaps onto the running child in
the rear display — the display's live view is sharp and legible in frame.
```
Both the active verb ("streams," "snaps") and the explicit physical path (origin → visible event → result) are important.
### Product hero shot
The hero scene treats the product like the best product cinematography does: clean composition, the named key light modeling its form, label exactly matching the reference image. Say it physically ("window light rakes across the body from camera-left, label centered toward camera") — never "glamorous product shot."
### Product placement
- Have the product already in hand or in the scene at video start — don't rely on the model to introduce it mid-video
- Describe specific placement and orientation rather than "the product is on the table"
---
## Product Fidelity — Match the Reference, Never Substitute
The product on screen must match the reference product image exactly in **type and form factor**, not just its label. The model will happily swap in a more common or generic variant of the category — over-ear headphones become earbuds, a flip-top bottle becomes a pump, a foil-wrapped bar becomes a boxed one. That is a total failure: when the product is shown, it must be *this* product.
- In every shot where the product is seen clearly — especially the hero scene and the closing frame — reinforce the match explicitly:
```
@reference_image[0] matches exactly — same device type, shape, proportions, controls/ports, and label.
Not a different model, not a generic stand-in of the same category.
```
- The product at the *end* must be the same product as at the start. A closing shot that shows a different-looking unit than the reference is one of the worst outcomes an ad can have — the payoff frame is the one the viewer remembers.
**Inspect the reference image and lock the actual part inventory.** The failure is not only "wrong model" — the model also **invents hardware the product does not have**, driven by the scene context. An earbud/headphone ad that mentions "calling" pulls the model toward an office-headset shape and it grows a **boom/stalk microphone** that isn't on the real product; a wireless product sprouts a cable; a clean earcup grows extra buttons. Look at what the reference image actually shows, then forbid the additions it is prone to:
```
The headphones are over-ear cups on a plain headband, exactly as in @reference_image[0].
They have NO boom microphone, NO mic stalk or arm, NO wire or cable, no extra buttons — add none of these.
Calls are handled by the built-in mics inside the earcups; nothing extends toward the mouth.
```
Add **no part that is not visible in the reference image**: no boom/stalk mic, wire, antenna, ear wings/hooks, strap, or extra controls. State the prominent absences as explicit negatives — "no boom mic, no wire" is what the model can act on; "true wireless headphones" alone is not. If a script action (like taking a call) would tempt the model to invent hardware, pin the absence right there in that scene.
---
## Concept & Usage Consistency — Nothing On Screen May Contradict What the Product Does
This is a top source of logically nonsensical generations. It has TWO parts — check both against `product_information`, not just the product name:
**1. Defining-feature contradictions** — nothing may contradict the feature the ad exists to sell:
- Open-ear / bone-conduction headphones → the ears stay **open**. No earbuds, in-ear tips, or over-ear cups on anyone in any scene.
- "No white cast" sunscreen → no chalky white residue left on skin after rub-in.
- Waterproof / sweatproof device → never shown failing or being protected from water.
**2. Usage-logic contradictions** — the character must interact with the product the way it is *actually used*, and must NOT simultaneously do the thing the product makes unnecessary. Read `product_information` for what the product *does* and what it *replaces*, then forbid the replaced behaviour:
- Earbuds/headset with built-in call mics → calls are handled hands-free **through the device**. Do NOT show the person holding a phone to their ear/mouth to talk — the phone stays in a pocket or held loosely for the screen; they speak while looking ahead, tapping the bud to answer.
- Wireless / true-wireless product → **no cables** running to a phone or player anywhere in frame.
- Hands-free / voice-controlled product → the hands are free for the activity, not occupied operating a substitute.
- A product that *is* the tool (a running watch that tracks pace) → the character doesn't also check a second device doing the same job.
The test for every human action in the script: *"Given what this product does, would a real owner still be doing this?"* If the product makes the action pointless or redundant, it is a usage contradiction — cut or replace it.
State the important contradictions as **explicit negatives** in the header (the CONSISTENCY line may hold more than one). "Open-ear design" alone does not stop the model from adding earbuds, and "great for calls" does not stop it from putting a phone to the ear; the negatives — "ears remain open, no earbuds on anyone" and "calls run through the earbuds, no phone held to the ear" — are what the model can actually act on.
**Absence-defined products need the rule REPEATED per scene — this is the one exception to state-once.** When a product is defined by *not* having a component the model strongly associates with the category (open-ear headphones = nothing in the ear; a cordless tool = no cord), stating the negative once in the header is not enough: the model re-adds the removed part scene by scene. In EVERY scene where the relevant body part is visible, add BOTH the positive and the negative, briefly:
```
Open-ear headphones, worn correctly: the band hooks over the ear and the pad rests on the cheekbone
in FRONT of the ear; both ear canals are visibly empty and open — no earbud, no ear tip, nothing inserted.
```
Describe the empty, open ear as a positive visible fact ("the ear canal is clearly empty and unobstructed"), not only as a prohibition — the model renders what you describe more reliably than what you forbid. Repeat this wherever a head or ear is in frame, every time.
---
## Product Mechanics — How It Opens, Sprays, Tears, Spreads
"Use the product" is not enough, and a generic pour path is not enough either. The model invents impossible mechanics: caps that hinge the wrong way, wrappers that dissolve, lotion that self-spreads and vanishes. Describe the *real* mechanism.
- **Opening** — name the closure and its exact motion:
- Flip-top cap → "the hinged cap flips UP and back on its hinge, staying attached to the bottle."
- Screw cap → "fingers twist the cap counter-clockwise, then lift it straight off."
- Foil bar wrapper → "fingers pinch the top seam and tear DOWN along the length; the foil splits at the seam and the bar slides part-way out."
- Pump → "the heel of the hand presses the pump head straight down."
- **Substance behavior** — real substances resist:
- Lotion/cream → "a thick bead sits on the skin and spreads only along the path the fingers actually rub, leaving a visible wet sheen — it does not self-spread, run, or vanish instantly."
- Liquid → a visible continuous stream from the opening to the landing point (see the pour example above).
- Never let a substance appear fully absorbed the instant it lands.
- **Consuming** — one unit stays one unit: "a single bar; one held in the hand, one bite taken and visible as a bite mark in that same bar." The count must not change between frames.
**When a manipulation renders badly, cut to the result instead of animating it.** Some fine hand manipulations — tearing a foil wrapper open, peeling a film lid, unscrewing a small cap — are actions current video models simply do poorly; forcing the full motion produces the "AI-like" morphing wrapper. Prefer to show the **result state** across a hard cut rather than the manipulation itself: instead of animating fingers tearing the foil, cut to the bar already cleanly opened — the foil neatly split along the top seam, the bar sitting part-way out, one hand holding it. If the opening must be on screen, keep it to a single simple motion in one shot (a clean tear straight down the seam), never a prolonged struggle, and never let the wrapper change shape, color, or count mid-tear. The wrapper's printed design stays fixed and matches the reference throughout.
---
## No Invented Text — The Garbled-Lettering Failure
Video models cannot spell. Any text it invents comes out garbled and misspelled: typo'd words ("work cmfortably"), misspelled wordmarks stamped on bags, banners, hurdles, or apparel, and nonsense background signage. Every one of those is an obvious AI tell and, for a brand, actively damaging.
- Do NOT ask the model to render words, signage, labels-with-copy, or logos-on-props anywhere in the scene. The **only** brand mark in frame is the one the reference product/logo already carries (on the actual product/packaging).
- Keep surfaces that would otherwise carry text plain — an unbranded gym towel, a blank banner, a clean wall — rather than inviting the model to letter them.
- All real on-screen text (a benefit super, the brand endcard) is rendered later as an **exact overlay graphic** and composited in. Never rely on the model to spell anything, including the brand name.
---
## Anatomy & Object Count — The Extra-Limb / Duplicate Failure
Close, busy, two-hand actions — rubbing lotion into an arm, tearing a wrapper, holding a device to the face — are the **highest-risk shots for extra or merged limbs and duplicated objects**. Constrain them explicitly:
- "Exactly two arms and two hands, belonging to the one person in frame — no third arm or hand, no duplicated or merged limbs, five fingers per hand."
- For hands-only application, name whose hands and how many: "her own two hands, nothing else enters the frame."
- Object count is explicit and constant: one bottle, one bar, one pair of headphones — never a second copy appearing mid-scene.
- Prefer framing that lowers the risk: a single hand demonstrating, or a locked clean composition, over a tangle of arms filling the frame.
---
## Common Failure Modes — Quick Reference
| Symptom | Cause | Fix |
|---|---|---|
| Over-saturated, HDR-look colors | Outcome color language / reference image color bleed | Broadcast-grade color block; restrict reference image to layout-only |
| Foreground looks processed, background natural | Differential rendering | Explicitly require consistent treatment across whole frame |
| Person looks airbrushed / plastic | Model's default face rendering | "Natural skin with visible texture and pores"; give them an action; or frame hands-only |
| Glowing halo/rim around subjects | Model's default "looks good" lighting | Allow backlight only from a named source; otherwise "no rim light, no edge glow" |
| Floaty, constantly drifting camera | Default model camera animation | Pick one support (locked/dolly/gimbal) and one motivated move max, or "camera locked, no movement" |
| Wrong objects in the scene | Reference image object copying | Explicitly ban copying; list exact items; end with "nothing else" |
| Showroom-sterile set | No explicit set dressing | List inhabited-world items explicitly with believable placement |
| Product action not visible | Vague action description | Describe physical path: origin → visible event → result |
| Wrong or generic location | Location not locked / stale reference image | Describe the chosen scene's concrete spatial anchors; ensure reference images match the chosen scene |
| Same scene appearing across different brands | Scene copied from an example or previous prompt | Re-derive the scene from this product's usage context; vary deliberately |
| Unwanted slow motion / speed ramps | Model emphasis default | "All action at natural speed"; slow motion only if scripted, with frame-rate cause |
| Wrong product type / generic substitute (earbuds instead of the over-ear model) | Only the label was pinned, not the form factor | "Matches the reference exactly — same type/form factor, not a different model" in the hero and closing shots |
| Invented hardware — a boom/stalk mic, wire, or extra buttons the product doesn't have | Scene context (e.g. "calling") pulls the model toward a different device shape | Inspect the reference image; state the absent parts as explicit negatives ("no boom mic, no wire") — pin it in the scene that tempts it |
| On-screen element contradicts the product's whole point (earbuds in an open-ear ad) | Concept guarded with a positive claim, not a negative | State the defining feature as an explicit negative in the header ("ears remain open — no earbuds on anyone") |
| Cap/wrapper opens impossibly; lotion self-spreads and vanishes | Opening mechanics and substance physics unspecified | Name the closure's exact motion and realistic substance behavior |
| Extra/merged arm or hand; duplicated product; wrong count | Model's default on busy close two-hand shots | "Exactly two hands, one person, five fingers"; explicit object count; simpler framing |
| Brand unclear until the end | No branded product visible in the opening | Brand bookend: the branded product is legible in the first scene and again in the last |
| Garbled/misspelled text or wordmarks on props, signage, apparel | Model asked to render invented text — it cannot spell | Ban invented text; only the reference logo carries a brand mark; real text is added as overlays |
| Prompt rejected (sensitive content) | Close-up face in reference image | Crop faces from reference images; describe people in text only |