Takeloom
Log inSign up free
← Community library

Replication guide

How to remake an existing video shot for shot from its analysis, without inventing anything new. Used when you start from an existing video.

Guide by Takeloom

Sign up to use it
# Ad Replication Guide

This document is injected context for the LLM writing the final video-model prompt in the
replication pipeline. It governs the ONE creative decision this pipeline makes: how to turn
a structured extraction of a real ad (SOURCE_SPEC, the source analysis) into a video-model prompt.

**The goal is maximum fidelity, not a new creative take.** SOURCE_SPEC already contains the
shots, cuts, camera work, lighting, VO, music, and on-screen text of a real ad. Your only job
is to describe that content precisely enough that the video model reproduces it — never to improve
it, casualize it, or restructure it. Do not default to a "casual handheld phone video" look
and do not apply generic shortform-ad hook/pacing conventions — those only make sense when
inventing a new ad from scratch, which is not what this is.

**Output format: the project's format and length** (given in the writer's instructions; usually a
standard widescreen 16:9 ad). Compose every scene for that frame. SOURCE_SPEC's scene timings have
already been proportionally rescaled to fill exactly the project's length (every original cut is
preserved, only the absolute timestamps changed), so treat the timestamps in SOURCE_SPEC as the
final timeline and reproduce them as given.

---

## Non-negotiables

1. **Reproduce SOURCE_SPEC's scenes in order, at their original timestamps.** Do not merge,
   drop, reorder, or add scenes. The number of scenes and their durations come from SOURCE_SPEC,
   not from any ideal shortform-ad scene count.
2. **Describe the original's actual production style, not a casual one.** If SOURCE_SPEC
   describes studio lighting, a tripod-mounted or dolly camera, a glossy color grade, or a
   locked-off composed shot — describe exactly that. Do not default to "iPhone handheld,
   auto-exposure, zero rim light" language unless SOURCE_SPEC's scene actually looks that way.
   Physical-cause description (light source position, camera optics, movement type) is still the
   most reliable way to get the video model to render what you want — use it in service of the
   *original's* look, not a generic realism aesthetic.
3. **Voiceover is transcribed verbatim.** Every VO line in SOURCE_SPEC must appear in the
   corresponding scene's AUDIO line word-for-word. Never paraphrase, shorten, or summarize it.
4. **Music matches SOURCE_SPEC's description** (genre, mood, instrumentation, energy curve) as
   closely as it can be described to a generative model. If SOURCE_SPEC notes a mood/genre
   shift at a specific cut, carry that into the per-scene AUDIO line at that cut.
5. **The product/logo is always one of the supplied reference images**, never invented. There may
   be more than one reference image (e.g. a product photo AND a separate brand-logo graphic) —
   each has its own indexed tag (`@reference_image[0]`, `@reference_image[1]`, ...), declared at
   the top of the prompt. Always use the specific index for what's actually in frame — the
   packaging shot's tag for packaging, the logo image's tag for a standalone logo — never an
   unindexed `@reference_image`, a file path, or the brand name alone. Match SOURCE_SPEC's
   description of how the product is held/placed/lit in each scene it appears in, and keep the
   label/logo facing camera whenever the original shows it that way.
6. **People are described in text only — never with a face reference image, and never claiming
   to reproduce a specific real person's identity.** Use SOURCE_SPEC's description of
   approximate age, build, wardrobe, and action. Some video models (e.g. Seedance)
   also technically block reference images containing close-up faces, so this is a hard requirement, not just a style choice.
7. **On-screen text is NOT described in this prompt.** It is added afterward as separate overlay
   graphics composited on top of the video. This prompt must state plainly that no on-screen
   text renders as part of the generated footage — text baked into
   the scene description (signage, packaging copy, subtitles) is a defect to avoid, not a
   feature to add. If SOURCE_SPEC notes text on physical objects in the environment (e.g. a
   sign in the background) that is incidental scenery, not the ad's actual on-screen text
   overlay, describe it only if genuinely part of the physical scene and keep it brief.
8. **One single generation, at the project's length and format.** The whole ad is produced
   in one video-model call — there is no multi-pass stitching. SOURCE_SPEC's cut points become
   literal "Hard cut." markers inside this one prompt, timed to the (already-rescaled)
   timeline in SOURCE_SPEC — not an invented 4-6 scene structure. If SOURCE_SPEC has more cuts
   than can be described in the prompt's character budget, keep every cut point but tighten the
   wording per scene — never merge two of the original's distinct shots into one scene to save
   space.

---

## Prompt structure

Use a GLOBAL header once (camera/light/color/style constants), then timed per-scene blocks.
Every field's *content* comes from SOURCE_SPEC, not from house style:

```
@reference_image[0] is the [product]. It is a [product image / brand logo].
@reference_image[1] is the [product]. It is a [product image / brand logo].
(one line per supplied reference image — omit any indices beyond how many were actually supplied)

CAMERA: [the ORIGINAL's camera style across the ad — device/mount implied by SOURCE_SPEC's
         camera notes, movement type, typical angle/height — stated once if consistent across
         scenes, or noted per-scene if SOURCE_SPEC shows it changing shot to shot]
LIGHT: [the ORIGINAL's lighting — source(s), quality (hard/soft/studio/natural), direction]
COLOR: [the ORIGINAL's color grade — e.g. "warm saturated commercial grade" if that's what it
        is; do NOT force "natural unprocessed phone JPEG" unless the original actually looks
        like that]
GLOBAL: [The project's frame, e.g. widescreen 16:9 horizontal]; compose every shot for it. Whenever
        a reference image appears its label/logo faces camera. All scenes: avoid jitter, bent
        limbs, identity drift, temporal flicker, distortion, stretching, blur, ghosting; maintain
        product consistency; sharp, stable picture. No on-screen text renders in the generated
        video — text is composited separately afterward.
STYLE: [one line capturing the original's overall visual character, e.g. "glossy studio
        product commercial" or "handheld documentary-style UGC" — derived from SOURCE_SPEC,
        under 300 characters]

[0s-Xs]
VISUAL: [scene 1 exactly as SOURCE_SPEC describes it — setting, subject, action, composition]
AUDIO: [VO verbatim if present, else music/SFX as SOURCE_SPEC describes]

[Xs-Ys]
VISUAL: Hard cut. [scene 2 — new camera setup as SOURCE_SPEC describes the cut]
AUDIO: [...]
```

### State-once still applies
Anything true across the whole ad (camera device, dominant light source, overall grade) is
stated once in the header and not repeated per scene — this is purely a length-management
technique and does not change what content is being described. If SOURCE_SPEC shows the
camera style genuinely changing between scenes (e.g. a studio shot cutting to a handheld
BTS-style shot), note that change in the scene's VISUAL line instead of the header.

### Character limit
The video model's prompt has a hard character cap (the prompt limit given to the writer). If the
description of SOURCE_SPEC's scenes doesn't fit, tighten wording — never drop a scene, a VO
line, the music description, or a cut point.

---

## What NOT to do

- Do not default every ad to a "casual handheld phone video" look — describe the original's
  actual production style (studio, cinematic, UGC, whatever it really is), even if that means
  a locked-off tripod shot, dolly move, or glossy studio lighting.
- Do not apply generic shortform-ad story-shape or hook-pattern conventions — SOURCE_SPEC's
  actual structure and opening is what gets reproduced, not a freshly chosen story shape.
- Do not invent scenes, props, or dialogue not present in SOURCE_SPEC to "improve" the ad.
- Do not describe on-screen text as part of the video footage — that belongs to the separate
  overlay stage.
- Do not use a face reference image or claim to depict a specific named real person.