← Community library
Vertical video guide
How vertical video works on Reels, TikTok and Shorts: the one-second hook, faces, safe zones, captions, pacing, direct address and loopable endings.
Guide by Takeloom
Sign up to use it# Vertical Video Guide — Reels, TikTok and Shorts (9:16)
This document is injected context for the LLM when writing and producing vertical short-form
video: Instagram Reels, TikTok and YouTube Shorts. It is the companion to the realism guide:
the **realism guide governs how each frame looks** (camera, light, color, people, no invented
text); **this guide governs how a vertical video works** (hook, framing, safe zones, captions,
pacing, direct address, the ending). Both apply at the same time.
**The core difference from TV.** A TV viewer has already chosen to watch; a feed viewer has not.
Every vertical video is watched on a phone held at arm's length, often without sound, by someone
whose thumb is already moving. The first second decides whether the rest exists. Polish matters
less than clarity, a face and a reason to keep watching — but "native to the feed" never means
"sloppy": the picture still follows the realism guide.
---
## The 5 Non-Negotiables
1. **A hook in the first second.** Not the first three — the first one. The opening frame already
shows something interesting and the first spoken word is already part of the hook. No logo, no
intro, no "hey guys", no slow reveal, no establishing shot.
2. **A face, large.** A person looking into the lens at close-up or medium close-up holds attention
better than anything else in the feed. Faces fill a big part of the frame, eyes in the upper
third.
3. **Readable on mute.** Most feed viewing starts silent. The hook and the key points must also be
on screen as text, inside the safe zones, and the picture itself must tell the viewer what is
going on.
4. **Something changes every 1–3 seconds.** A cut, a new angle, a new object, a gesture, a text beat.
Stillness longer than about three seconds reads as the end of the video.
5. **An ending that pays off — or loops.** Deliver what the hook promised, then either end on a
clear call to action or cut back so the last frame flows into the first.
---
## The Frame — 9:16 and Its Safe Zones
The canvas is 1080x1920 (or 720x1280). The platform interface covers large parts of it:
| Zone | Covered by | Rule |
|---|---|---|
| Top ~14% (about 270 px of 1920) | Status bar, account name, "Following / For you" tabs, close button | No text, no faces' eyes, no key action |
| Bottom ~20% (about 380 px of 1920) | Caption, username, audio track, progress bar, comment box | No text, no product, no key action |
| Right edge ~12% (about 130 px), lower half | Like, comment, share, save and profile buttons | Nothing important in the lower-right |
| Left edge ~4% | Swipe margin | Keep text off the edge |
**The safe centre** is roughly the band from 14% to 80% of the height, slightly left of centre.
Everything that must be seen — the face, the product, the hook text, the key points — lives there.
- **Faces:** eyes at about one third from the top of the safe band (roughly 25–30% of the full
height), never in the top 14%.
- **Hook text:** top of the safe band, centred, large (it doubles as the thumbnail title).
- **Captions:** the lower part of the safe band, around 60–75% of the height — above the platform
caption, never at the very bottom.
- **The product or object:** held at chest height, which in a close framing lands in the middle of
the frame.
- Design for all three platforms at once: if it is safe on the most crowded one (TikTok), it is
safe on all of them.
---
## The Hook (0–1 s)
Pick one pattern and commit. The hook is visual and verbal at the same moment.
- **The result first.** Show the finished thing, then explain how ("This took four minutes.").
- **The bold claim.** A specific, slightly surprising statement said straight into the lens
("You're storing bread wrong.").
- **The question.** One the viewer already has ("Why does your rice always stick?").
- **Mid-action.** Join a moment already in motion — a pour, a jump, a door opening — with the first
word landing on the first frame.
- **The pattern break.** An unexpected image or angle: an extreme close-up, an object moving
toward the lens, a presenter entering frame already talking.
- **The list promise.** "Three things I wish I knew before…" — tells the viewer exactly what they
get and how long it takes.
Rules:
- Write the hook line verbatim and put the same words (or a shorter version, at most 6 words) on
screen from 0 s.
- The first frame must work as a still image: it is the thumbnail in grids and previews.
- Never spend the hook on the brand name, a greeting or context. Context comes after the viewer has
decided to stay.
---
## The Presenter — Direct Address
A presenter talking to camera is the native language of the feed.
- **Eyes into the lens**, at eye level, phone-distance: medium close-up, 0.6–1 m, a 24–35 mm
equivalent lens. It should feel like a video call from a friend who knows something.
- **Speak like a person, not an announcer.** Short sentences, contractions, one idea per sentence,
the words a real creator would say. About 2.5–3 words per second: faster than TV, never rushed.
- **Energy comes from specifics, not volume.** Gestures that show (holding the object up, counting on
fingers, pointing at the thing) beat shouting.
- **Match voice and face.** A visible speaker's voice matches their apparent gender and age, and the
lips move with the words. Use a separate off-screen voice-over only for cutaways without the
presenter.
- **Same person every shot.** Face, hair, wardrobe and setting are described identically in every
scene; the presenter is the continuity anchor of the whole video.
- Cutaways (hands doing the thing, the object close up, the result) carry the voice-over over them
and return to the face at least every 5–6 seconds.
---
## Captions and On-Screen Text
- **Captions for spoken words** are standard in vertical video: short chunks of 2–5 words, in time
with the speech, high contrast (white with a dark outline or a solid box), in the lower middle of
the safe band. They are added in the editor as exact text layers — never asked of the video model,
which cannot spell.
- **Key-point text** is separate from captions: the hook line, "Step 2", a number, a claim. At most
6 words, large, top or centre of the safe band, on screen for at least 1.5 s.
- One font family for the whole video; one accent colour at most.
- Never put text in the top 14%, the bottom 20% or behind the right-hand buttons.
- The words on screen and the words spoken tell the same story; together they must still make sense
with the sound off.
---
## Pacing and Structure
- **Scene length:** 1–3 seconds for cutaways, up to 4–5 seconds for a presenter shot that is
delivering a line. A 15-second video usually has 5–8 shots; a 30-second one 10–15.
- **Structure for 8–15 s:** hook (0–1 s) → one development or demonstration (1–12 s) → payoff and
call to action or loop (last 2–3 s).
- **Structure for 30–60 s:** hook → promise ("here are three…") → points, each with its own visual
change → payoff → call to action. Re-hook at the midpoint with a new visual or a "but the last one
matters most".
- **Cuts on words.** Cut on the beat of the speech or the music, not in the middle of a word.
- **Jump cuts are fine** for a presenter: removing pauses between sentences keeps energy, as long as
the framing changes a little (a punch-in from medium close-up to close-up) so it reads as a choice.
- **Something new in every shot:** a new angle, a new object, a gesture, a text beat. Two adjacent
shots that could be one take are one shot too many.
- **Music:** a bed that fits the platform's tone, lower under speech, with the energy changes timed
to the cuts. Trending sounds are added by the creator on the platform, not generated.
---
## Endings — Payoff, Call to Action, Loop
- **Deliver the promise.** The last seconds show the result, the final point or the answer to the
opening question — never just trail off.
- **Call to action:** one, specific and cheap to do ("Follow for part two", "Save this for later",
"Try it tonight"). Spoken and on screen for at least 2 seconds, inside the safe centre.
- **Loopable endings** raise watch time: end on a frame and a sentence that lead straight back into
the first ("…and that's why you should never—" / "You're storing bread wrong."). Match the
framing of the last shot to the first so the loop is seamless.
- Do not end on a logo card or a black frame. If there is a brand, it appears in the scene (the
product in hand, the logo on the pack), not as an end slate.
---
## Platform Notes
| | Instagram Reels | TikTok | YouTube Shorts |
|---|---|---|---|
| Length that works | 7–30 s; up to 90 s for tutorials | 9–30 s; longer for stories and tutorials | 15–45 s; up to 60 s |
| Tone | Polished, aesthetic, aspirational; lifestyle and how-to | Raw, personal, fast, humorous; creator voice | Informative, curiosity-driven; explainer and list formats |
| Hook | Visual first, beautiful opening frame | Verbal first, a line that starts mid-thought | A question or a promise; the title and the first line work together |
| Text | Clean, minimal, one font | Native captions are expected | Clear key-point text; the title matters for search |
| Ending | Save-worthy value, soft call to action | Loops and "part two" | Subscribe prompts work; loops also count views |
When the brief says "All", design for TikTok's safe zones and pace, keep the look clean enough for
Reels and make the first line clear enough for Shorts.
---
## Realism in Vertical
The realism guide still governs every frame. Vertical-specific notes:
- **Phone-native does not mean amateur.** A creator look is a stable handheld or a small tripod at
eye level, soft window light, a real room — not shaky footage and not a studio.
- **Vertical composition:** stack things top-to-bottom instead of side-to-side; a two-person shot
places the people one slightly behind the other, not at the edges of the frame.
- **No text in the world:** the video model still cannot spell. Signs, screens and packaging copy
stay blank; all words are overlays.
- **Hands close to the lens** are high-risk for extra fingers: state exactly two hands, five fingers
each, and keep hand demonstrations simple.
---
## Self-Check Before Submitting
- Does the very first frame, as a still, make someone stop scrolling? Is the first word part of the
hook?
- With the sound off, do the picture and the on-screen text tell the story?
- Is every face, product and word inside the safe centre — nothing in the top 14%, the bottom 20% or
behind the right-hand buttons?
- Does something change at least every three seconds?
- Does the presenter look into the lens and sound like a person?
- Does the ending pay off the hook, with one clear call to action or a seamless loop?