← Help Center

Every tool

The rest of the tools

Silences, filler words, bad takes and captions have guides of their own. These six are the rest of what is in the panel today. Each section says what the tool changes, because that is the thing worth knowing before you press anything.

All of them work the same way: choose a source, scan, review what came back, then approve. Nothing reaches Premiere until you approve it, and every tool card in the panel says what kind of change it makes — Writes keyframes, Creates sequences, Markers only, Review only.

Add punch-ins

Changes: Motion Scale keyframes on clips already in your timeline. No cuts.

A punch-in is a small push toward the speaker on a line that matters. Um Out finds the moments and writes the keyframes; you pick how far it pushes.

One control: Zoom strengthSubtle (105%), Balanced (110%) or Bold (115%). Anything past about 115% starts to show softness in the image, which is why the slider stops there.

Add smart zooms

Changes: Motion Scale keyframes. No cuts.

The same mechanism as punch-ins, but you choose what triggers them rather than letting Um Out decide:

  • Voice changes — when speaking returns after a pause.
  • Edit points — at existing V1 timeline cuts.
  • Emotional turns — at expressive questions and reactions.
  • Key ideas — on meaningful statements, which needs an AI provider.

Motion feel sets the curve: Direct is an instant scale change, Glide eases in and out, Quick push has a fast opening ramp. Zooms land at 107% and never closer together than 1.5 seconds, so a talkative stretch does not turn into a slideshow.

Find B-roll

Changes: inserts clips on V2 with their audio muted.

Um Out reads your transcript, then looks for footage already in your project bin whose filename matches what you were talking about. Approved shots land on V2; removing one is a single delete.

It matches on filenames, so your filenames are the feature. drone-city-02.mp4 matches when you talk about the city. DJI_0043.mp4 matches nothing, ever. Renaming your B-roll is the single fastest way to better results. Um Out does not search stock libraries and does not look at the pictures — if the name says nothing, the clip is skipped rather than guessed at.

Build short clips

Changes: nothing. It creates new sequences and leaves your timeline alone.

Each approved moment becomes its own 1080 × 1920 sequence named “Um Out Short”, 15 to 45 seconds long, with captions rendered in whatever look you set on the Captions tool.

Moments are ranked on three things, and the panel shows you the reasoning for each: the opening (short-form is decided in the first two seconds), the reaction (how hard the room responded, read from the waveform), and the payoff (whether something actually lands in the back half).

Two honest limits. It does not reframe a moving subject — the portrait crop is fixed, so a speaker who walks out of frame walks out of frame. And “viral” here means a shortlist worth watching, not a forecast; nothing in it models a thumbnail, a title, or what an algorithm wants this week.

It needs picture on V1 and refuses without it, before transcribing rather than after.

Clip That

Changes: nothing. Each call-out becomes its own sequence.

For anyone who says “clip that” while recording. Um Out finds every one of those call-outs in the transcript and cuts the moment around it.

  • Seconds before — 30 by default, because the moment is already over by the time you say it. A moment that took two minutes to build needs this raised.
  • Seconds after — 2 by default. Zero cuts the call-out itself out of the clip.
  • Shape — vertical 9:16 or landscape 16:9.
  • Your own phrases, one per line, with an option to ignore the built-in list entirely.

It matches whole phrases, so “clip” on its own never fires and talking about a paperclip is safe. It also only finds what was said out loud and transcribed — a call-out muttered under a loud moment leaves nothing to match, and a missed one looks exactly like no call-out at all.

Mark chapters

Changes: markers only. This is the one tool that cannot alter your edit.

A model reads the transcript and marks where the subject changes. Not where the pauses are — a long pause is not a chapter, and the sentence where you stop describing the problem and start describing the fix usually has no pause in front of it.

With no AI provider configured it falls back to splitting on the longest pauses and naming sections after the words that follow. The results screen tells you when that happened rather than passing a mechanical split off as understanding.

Chapters land at least 45 seconds apart, up to 30 of them, and the first is always pulled to 0:00 — YouTube silently rejects a chapter list that starts later.

Find stream highlights

Changes: nothing, ever. Review only.

This one has its own guide.


Something behaving unexpectedly? See known issues and workarounds, or the troubleshooting guide.