Add B-Roll to a Talking-Head Video Without Stock Footage
The problem with b-roll for talking-head video isn't finding footage — it's that the footage you need doesn't exist. No stock library has a clip for "our churn dropped from 6% to 2.4% after we changed onboarding." You end up laying down a shot of someone typing on a laptop and hoping nobody notices it has nothing to do with the sentence. This guide covers the alternative: cutaways generated from what you actually said, and a way to add visuals without cutting away at all.
Why stock fails talking-head video
Talking-head content is carried by what you're saying. The moments that need a visual are almost always the specific ones — a figure, a comparison, a process, a term. Those are exactly the moments a stock library can't serve, because stock is generic by design.
The result is a familiar compromise: cutaways that are decorative rather than explanatory. They fill the frame, they hide a jump cut, and they carry no information. Viewers don't consciously notice — but the video ends up feeling padded rather than dense.
Four ways to get b-roll that isn't stock
- Shoot it yourself. Best quality, highest cost. Practical for product shots, not for abstract points.
- Screen recordings. Free and specific — ideal when you're talking about software.
- AI video generation. Good for atmosphere and impossible shots; expensive per second, and still can't render a number correctly.
- Generated motion graphics. Designed frames built from your script: numbers, charts, comparisons, kinetic type. The right tool when the point is information rather than imagery.
Most good talking-head videos mix at least two of these. This guide focuses on the last one, because it's the one that scales without a shoot day.
Generating cutaways from your own words
Your edit already contains the source material: the subtitles. An .srt file is
your words plus the exact time each line starts and ends — which is everything needed to
author a matching visual per line and place it correctly.
- Export subtitles from your edit (CapCut, Premiere, Resolve and YouTube all export
.srt), or upload the audio and let it be transcribed. - Generate the track. Each line becomes a designed frame; the whole thing comes back as one continuous video the length of your edit.
- Lay it on the timeline. Because the timings came from your subtitles, it lines up with the cut you already made — there's nothing to nudge.
- Cut holes where you want your face. Use the generated track as the base layer and trim it wherever you'd rather be on camera.
The full walkthrough is in adding b-roll automatically from subtitles, and the transcript-first version is in turning a transcript into motion graphics.
Staying on camera: overlay graphics
Cutting away has a cost: you lose the face. For a lot of talking-head content — personal brands, opinion, teaching — the face is the retention mechanism. Cutting to a chart for eight seconds can lose more than the chart gains.
The alternative is to add the graphic on top of yourself. Overlay mode returns a transparent-background video — ProRes 4444 with a real alpha channel, not a black rectangle — containing only the graphics. You drop it on the layer above your footage and stay on camera the entire time: a number appears beside you as you say it, a lower third names the term, a small card holds the comparison.
Practically, this is the same workflow as a news broadcast: the presenter never leaves the frame, and information arrives around them.
Keeping graphics off your face
An overlay is only useful if it doesn't cover the thing it's decorating. Two rules do most of the work:
- Declare where you are. Drag a box over the region to keep clear — your head and torso — and every element gets placed outside it. Guessing where a person is in frame is the single most common way overlays go wrong, so it's worth the five seconds.
- Leave the caption band alone. The bottom sixth of the frame belongs to subtitles on every short-form platform. Graphics that drift into it get covered by burned-in captions or platform UI.
Vertical video changes the layout
On a 16:9 canvas a person occupies the middle third, so graphics can live in a left or right column. On 9:16 that column doesn't exist — a person fills most of the width. The workable space is two horizontal bands: a strip along the top, and a lower third that sits above the caption area.
This is why the aspect ratio has to be chosen before rendering rather than cropped afterwards: the composition rules are genuinely different, and a 16:9 track squeezed into 9:16 puts text straight across your face.
How much is too much
The usual advice is that cutaways should support the A-roll rather than replace it — for most talking-head content that lands somewhere between a quarter and a third of runtime. Overlay graphics sit outside that budget, because they add information without taking you off screen; the constraint there is visual noise instead. A useful discipline: at most one or two elements visible at once, and every element earns its place by carrying a fact.
Frequently asked questions
How do I add b-roll to a talking-head video without using stock footage?
Can I add graphics without cutting away from my face?
Will the graphics cover my face?
How much b-roll should a talking-head video have?
Does this work for vertical short-form video?
What if I already edited the video?
Related: Transcript to motion graphics · Add b-roll automatically from subtitles · What is b-roll? A-roll vs b-roll
More from Guides · See pricing or read the Privacy Policy.