A weekly podcast channel is really two channels. There is the 40-minute episode, which wants chapters, tight silences and clean audio. And there are the five or six vertical clips, which want captions, punch-ins and a hook in the first second. Most editors treat these as separate jobs, which is why the clips get made on Sunday night or not at all.
They do not have to be separate. Both jobs read from the same source of truth — what was said, and when — and if you only build that once, the second job needs no second transcription, no re-upload, and no re-scrubbing to find the same moments again.
Step 1: transcribe before you cut anything
This inverts the habit of cutting first and captioning later, and it is the whole trick. A transcript of the raw assembly is a searchable map of the episode: you find the good bits by reading rather than by scrubbing back and forth through 40 minutes of audio looking for a line you half-remember.
Transcribe the full assembly, not the finished cut. You want the transcript to cover material you might still use.
- With the raw assembly sequence open, go to Window → Text → Transcript, then Transcribe.
- Premiere hands back a scrollable, timestamped transcript in the Text panel — click any line and the playhead jumps straight to it.
- Read it top to bottom instead of scrubbing the timeline. Note the timecode of every section break and every moment you'd clip on their own — you'll turn both into markers in step 3, off this same read.
Step 2: pull the silences out of the long-form cut
Two people talking generates a lot of dead air — thinking pauses, overlaps, the half-second before someone answers. Trimming those is what makes a conversational edit feel tight.
By hand, this is ear work, one pause at a time. On a 40-minute two-person conversation that is a lot of marking — budget well over an hour if you are cutting every pause this way.
- Play the raw assembly at normal speed and listen for any pause longer than half a second.
- Press C for the Razor tool, then click at the start and the end of the pause to cut it free from the clips either side.
- Select the isolated gap and press Shift+Delete, or right-click it and choose Ripple Delete, so the audio either side snaps together.
- Repeat through the full episode, then scrub back at double speed once to catch any cut that landed mid-word.
Be conservative on the first pass — this is ear judgment, not a threshold slider, and it's easy to over-cut once you're in the rhythm of it. Cutting every pause to zero makes a conversation sound like an argument; leaving a beat of air is what keeps it sounding like people.
Step 3: mark the clips while you are already reading
You are in the transcript picking chapter points anyway. That is the cheapest possible moment to also mark the six moments worth clipping, because you are already holding the whole episode in your head. Drop a marker on each one as you go, and name it while the line is still in front of you — a marker named "silence" tells you nothing in twenty minutes, a marker named "the pricing objection" does.
- Press M at every topic change for a chapter marker, and again at every moment worth its own clip — same read, two different marker sets.
- Double-click each marker in the Timeline panel and type a real name into the Name field before you move on to the next line.
- Right-click a marker to give it a colour — one colour for chapters, a different one for clip candidates — so the two sets are easy to tell apart once you're done reading.
- Open Window → Markers to see every marker in one column with its timecode, sortable and filterable by colour.
A moment worth clipping usually starts one sentence earlier than you think. The setup line is what makes the payoff land for someone who arrived cold.
Chapters come out of the same pass, but they are not automatic once marked: read the timecodes off the Markers panel and type them into the YouTube description as 0:00 Title, one line per marker, with the first line sitting at 0:00 — that's the format YouTube turns into clickable chapters. The clip markers don't go in the description at all; they're your map for step 4, nothing more.
Step 4: build the verticals off the same transcript
Caption the full episode before you duplicate anything, not after. Premiere hands every duplicated sequence a new internal sequence ID, and Backstage Cut's transcript cache is keyed to that ID — so generating captions on a sequence you've just duplicated is a fresh transcription, not a free one. Generate them once on the long-form sequence instead: they land as ordinary clips on their own track, and duplicating that sequence per clip carries those clips along with everything else already sitting on the timeline. (If you don't want word-timed captions in the long-form deliverable itself, hide or delete that track afterward — deleting it from the original doesn't touch the copies already sitting in the duplicates you've made.)
- With the long-form sequence still the original (not yet duplicated), generate captions for the whole episode.
- Duplicate the sequence and set the frame size to 1080×1920.
- Trim to the marked moment, starting one sentence before the payoff — the caption clips in that range come along with the trim.
- Reframe the speaker: in Effect Controls, twirl down Motion, raise Scale until the frame is filled, then set Position so whoever's talking sits centred in the crop — Scale alone doesn't recentre anyone who wasn't already in the middle of the original 16:9 frame. On a two-person shot, keyframe Position again at each point the mic switches, so the crop follows whoever's actually talking instead of leaving them off to one side. Reposition the caption clips to sit inside the vertical safe area once the crop is set — they were placed for the original 16:9 frame.
- Add punch-ins on the emphatic lines — a second Scale keyframe pair a few frames apart, eased in and out rather than linear — and cut on the beat where the other person reacts.
Repeat for each marker. By the third clip this takes minutes, because every decision that required judgement was made in step 3.
The mistakes that undo the one-pass idea
- Trimming a clip's in-point exactly on the payoff line instead of a sentence earlier. A moment worth clipping usually needs its setup to land for someone who arrived cold — cut on the line itself and the clip is technically correct and makes no sense.
- Scaling the reframe without keyframing Position on a two-person shot. Raising Scale alone recrops whoever wasn't already centred in the original 16:9 frame toward one edge instead of onto them.
- Naming markers "good part" or "clip 3" instead of the line itself. That name is meaningless three weeks later when you're finally back to cut the clips, and the whole point of marking during the read was not having to re-find the moment by ear.
- Using one marker colour for both chapters and clip candidates. The two sets blur together in the Markers panel, and you spend the clip-cutting session re-reading which marker is which instead of cutting.
- Exporting all six clips before the host signs off on the long-form episode. A late trim to the episode after that point means re-cutting every vertical clip against the new timestamps, not just retrimming the duplicated sequences you already built.
- Duplicating the sequence first and generating captions on the copy. The duplicate gets its own sequence ID, so Backstage Cut treats it as unseen footage and re-transcribes — caption the original before you duplicate, not after.
Why this is the shape Backstage Cut is built around
Backstage Cut transcribes the active sequence once and caches it, then drives captions, chapters, silence cutting and transcript-timed zooms off that one transcription. The 40-minute episode and the six clips cut out of it run from a single transcribed pass — you are not billed minutes again for footage you have already transcribed.
Everything it produces is ordinary Premiere clips and keyframes you can still retime or delete by hand, which matters more here than anywhere else: a podcast edit changes after the clips are marked, and a workflow that locks you out of a late change is a workflow you throw away when the host asks for one more trim.
The habit that actually saves the time
None of the above depends on a particular tool. The transferable part is the order: transcribe first, read once, mark everything you will need in that single read, and only then start cutting. Editors who do the clips on Sunday night are usually editors who read the episode twice.