Guides · Workflow

Keep the take where the explanation actually landed, cut the rest

A tutorial or course video is rarely one clean take. It's a screen recording, a separate webcam, and an explanation you said three different ways before one of them came out right — and all three usually get discovered in that order, at the worst possible point in the edit. Here's a pass order that catches the retake before it costs you a callout you'll only have to redo.

· Updated October 1, 2026 · 8 min read

Why a tutorial edit breaks differently than a talking-head one

A talking-head video runs on one camera and forgives a rough sentence — a viewer reads intent, not a transcript. A tutorial doesn't get that forgiveness. Get a menu path wrong, or leave a callout pointing at the shortcut you meant to say and didn't, and it reads as broken software, not an editing slip. That accuracy requirement is the whole difference, and it changes what order the edit should happen in.

It also usually runs on two sources instead of one — a screen recording and a separate webcam, captured by different apps and rarely started in the same frame — and on explanations that get restarted mid-sentence more than a normal talking-head take does, because getting a technical step exactly right takes more tries than a casual aside does. Building callouts, zooms and captions before those two problems are resolved means redoing all of it once they are.

Four passes for a screen-recording tutorial

Pass 1: Sync the sources and pick your keeper take

Skip the syncing half if your screen recorder already captured the webcam in the same file — plenty of course creators record that way, and there's nothing to line up. This pass matters when the webcam was shot separately, on its own camera or app, for better quality than a recorder's built-in overlay.

  1. Import the screen recording and the webcam clip into the project, both untrimmed.
  2. Select both clips in the Project panel, right-click, and choose Merge Clips.
  3. Set Synchronize Point to Audio if the webcam clip picked up any version of your own voice — even faintly through the room mic while you spoke into a separate mic for the screen recording — and click OK. Premiere reads both waveforms and lines the video up on its own.
  4. No usable audio on the webcam track at all? Set Synchronize Point to In/Out Points instead, and start both recordings on the same visible countdown next time — that's the only reliable manual sync point when there's no audio to match.
  5. Drag the merged clip to the timeline. Both sources now move together as one clip, so a trim on one moves the other with it — which matters for the second half of this pass.

With the sources synced, the second job in this pass is deciding which attempt at each explanation actually stays.

  1. Play through at normal speed and press M at every point where you stopped mid-explanation and started over — you're marking retakes here, not judging them yet.
  2. For each marked restart, decide which attempt actually explains it best. That's usually the last one, but not automatically — sometimes the final take rushes just to get it over with, and an earlier attempt is the clearer one.
  3. Working from the last marked retake in the sequence backward to the first, press C for the Razor tool, cut the attempt you're not keeping free at both ends, select the isolated range, and Shift+Delete to ripple-delete it, closing the gap on every linked track at once. Deleting front-to-back shifts every marker and timecode after the first cut, so cutting from the end keeps the retakes you haven't dealt with yet exactly where you marked them.

Pass 2: Read the transcript for what needs a callout or a zoom

With retakes cut and the sequence at its real length, transcribe it and read the page rather than re-watching. A keyboard shortcut or a menu path is far easier to catch reading than by ear, especially the ones you say quickly because your fingers already know them.

  1. Window → Text → Transcript, then Transcribe, on the now-trimmed sequence.
  2. Read the transcript top to bottom in the Text panel rather than scrubbing the timeline.
  3. Press M at every moment that needs an on-screen callout, and type the condensed version straight into the marker name — the exact key combo or menu path, not the full sentence you said.
  4. Mark a second, separate set of moments: anywhere the thing you're pointing at is too small to actually read once compressed for upload — a checkbox, a tab, an icon in a toolbar. These need a punch-in, not a callout, and mixing the two lists together slows pass 3 down rather than speeding it up.

Pass 3: Frame the webcam once, zoom the UI where it needs it

The webcam bubble only needs building once, then Paste Attributes carries it onto every other segment.

  1. Select the webcam clip, open Effect Controls (Window → Effect Controls), and apply the Crop effect. Trim Left, Right, Top and Bottom until the frame is square — most webcams shoot 16:9, and an ellipse mask is only ever as round as the frame it's drawn on.
  2. Still in Effect Controls, twirl down Opacity and click the ellipse icon to Create Ellipse Mask (or the rectangle icon for a rounded rectangle instead). Hold Shift while dragging a corner handle so it resizes as a circle rather than stretching into an oval.
  3. Twirl down Motion and set Scale and Position so the bubble sits in whichever corner is actually clear — check it against a frame where your screen-recording app's own toolbar or cursor is visible, not an empty one, since that's where it tends to collide.
  4. Select the webcam clip, Edit → Copy, then select every other webcam segment in the sequence and Edit → Paste Attributes so the same crop, mask and position land on all of them at once instead of rebuilding it per clip.

The zoom-on-tiny-UI half of this pass is the part specific to a screen recording, and it has to be built per marker rather than copied once, since the target sits somewhere different on the screen each time.

  1. For each marker from pass 2's zoom list, select the screen-recording clip and open Effect Controls (Window → Effect Controls).
  2. Set a keyframe on Scale and Position a few frames before the marker, at the untouched values.
  3. Move the playhead to the marker and keyframe Scale up — 150 to 200% is usually enough to make an 11-point menu label readable at 1080p — dragging Position so the target sits centred in frame.
  4. Hold that framing for at least a second and a half past where you'd instinctively pull back out. A zoom already retreating by the time a viewer's eye arrives at the target defeats the point of zooming in at all.
  5. Keyframe Scale and Position back down to the original values a beat after you move on verbally, not the instant you do, so the release matches the sentence rather than a fixed duration.

This is close to what AutoZoom does on a talking-head video, aimed at a different target — there it's landing a punch-in on the line that carries the beat; here it's making sure one specific pixel is legible. That distinction is also why this pass stays manual regardless of what else in the edit is automated: reading a screen recording's own UI for a legibility problem isn't the job AutoZoom's transcript-timed punch-ins are built for.

Pass 4: Callouts, captions, chapters, one full watch at real size

Only now, with retakes cut and the zooms built, add the layers that sit on top of the locked picture. If pass 1's retake cuts moved anything meaningful, re-transcribe before generating captions off it — a caption timed against a timestamp that no longer exists drifts the moment someone plays the video back.

  1. Build the first callout in Essential Graphics (Window → Essential Graphics, New Layer → Text), style it once, then copy it onto every other marked moment and swap in that marker's text.
  2. Generate captions from the current transcript — re-transcribed if pass 1 moved anything.
  3. Mark chapters off the same transcript: at least three, the first at 0:00, none closer than ten seconds apart, pasted into the description as 0:00 Title Here.
  4. Watch the whole thing once at the sequence's actual export resolution, not a shrunk Program monitor. A zoom that reads fine at half-height can still land on an unreadable label once compressed for upload, and full-size playback is the only pass that catches it.

The mistakes that make a tutorial hard to follow

  • Building callouts or zooms against a retake that gets cut later in the same session, so the timing work has to be redone against whatever replaces it.
  • A webcam bubble that drifts to a different corner or size between clips because each one was framed by eye instead of Paste Attributes from the first.
  • A punch-in that retreats before a viewer's eye has actually reached the thing it zoomed in on.
  • Captions generated before a mid-edit script change, so the words on screen and the words spoken quietly disagree for the rest of the video.
  • No chapters on anything past a few minutes, so a viewer looking for one specific step has to scrub the whole recording to find it.

Where Backstage Cut fits into passes 1 through 4

Backstage Cut is still actively developed, and the pass order above is the actual time-saver — it holds whether or not you install anything. Where the panel helps a screen-recording tutorial specifically is Retake Cut, in pass 1: it reads the transcript for lines said more than once, catching a second attempt even when it came out worded differently rather than matching strings, and lists every take with the one you likely meant already selected — most of the keeper-picking above, done from a list instead of a full re-watch.

A-Roll PiP frames the webcam bubble once, against your own footage, and applies it to every clip on that track, or only the ones you select — pass 3's framing half. Emphasis Text finds the numbers, shortcuts, UI paths and warnings itself and pops them on the exact word they were said on, which covers most of pass 2 and pass 4's callout work off the same transcript captions and chapters already run from. None of that reaches the tiny-UI zoom in pass 3 — that stays a manual keyframe pass, the same as it is above.

The habit that actually makes a tutorial repeatable

None of the four passes above require anything beyond Premiere, a screen recorder and a webcam. What they require is not skipping the order: settling on a keeper take before styling a single callout, reading the transcript instead of re-watching for what needs a zoom, and checking the final result at real size instead of trusting a shrunk Program monitor. An editor who works this way isn't faster at explaining a feature than one who doesn't — they're just never re-cutting a callout against a take that already got deleted.

Try it on the video
you're editing tonight

7 days of unlimited transcription is plenty for a whole video, start to finish, on your own footage — no card, and nothing to undo if you don't like it.

50% off — join the community