Home
Dubb

Micro-Clipping: How to Find the Good Clips in a Two-Hour Video

Ruben

Ruben

Micro-clipping works from the transcript rather than the footage. A subtitle file carries every spoken line with its timecode, so an AI reading that file can name the start and stop points of a few strong segments. You then cut only those, instead of reviewing a pile of automatically generated candidates to find the ones worth keeping.

What Is Micro-Clipping?

Micro-clipping is the practice of extracting short, self-contained segments from a long recording so they can be published on their own. It differs from editing in that nothing is being assembled: the clip already exists inside the source, and the entire job is deciding where it starts and where it ends.

The approach here uses a subtitle file as the working material. Dubb Agent has a micro-clipping skill that reads the transcript and returns timecoded segments, which you then cut in an editor. The source recording can be a webinar, a podcast, a training, a demo, or anything else long enough that watching it back is not realistic.

Table of Contents

  1. Candidates Versus Precision
  2. What You Actually Need: A Transcript File
  3. Where to Get an SRT or VTT
  4. What Is Actually Inside the File
  5. The Micro-Clipping Skill
  6. The Prompt
  7. Judging What Comes Back
  8. Cutting the Clip
  9. Removing the Silences
  10. Format and Length
  11. Clipping Approaches Compared
  12. Common Mistakes and Troubleshooting
  13. Best Practices I Actually Follow
  14. Proof: Why This Actually Works
  15. Frequently Asked Questions
  16. Three Clips Is Enough

Candidates Versus Precision

Micro-clipping has become its own software category, sometimes written as micro clipping, and the tools in it mostly work the same way: feed in a long video, get back a set of suggested clips, pick the good ones.

That design optimises for recall. The goal is to make sure the best moment is somewhere in the pile, which means generating generously and letting you sort it out. It is a reasonable choice and it is why these tools rarely miss anything.

The cost lands somewhere people do not count when they estimate the time saved. You have not removed the work, you have changed its shape: instead of editing you are now reviewing, and reviewing near-misses is slower and duller than it sounds. Working through fifteen candidates to find two you would actually publish is an afternoon, and it is an afternoon of judging rather than making.

We used OpusClip for a while and the editing quality genuinely impressed us. What did not work for us was the review step. On our material the proportion of candidates worth publishing was low enough that the sorting ate the saving, and that is a fit problem rather than a quality problem.

Asking for three specific clips from a transcript optimises for the opposite thing. You will miss good moments, because you asked for three and the video contained six. In exchange you get three you can cut immediately. For a daily publishing habit, three usable clips beats fifteen maybes, and that is the whole argument.

What You Actually Need: A Transcript File

Micro-clipping has one requirement, and it is the thing that makes the rest possible: a subtitle file for your recording. Either an SRT or a VTT. They are close enough to interchangeable for this purpose.

The reason this matters is that a model cannot watch your video, but it can read every word you said along with exactly when you said it. That turns a two-hour recording into a document, and finding the three strongest passages in a document is something a model is genuinely good at.

No transcript, no AI-based clipping. That is the gate.

Where to Get an SRT or VTT

Several routes to the file micro-clipping needs, and you probably already have one available.

From your video platform. Upload the recording to Dubb Studio and captions are generated for it, after which you can download the SRT or the plain text. Captions need to be generated for that video, so on an older recording you may need to trigger it before the download option appears. Transcription is a paid-plan feature, and generating captions does not consume additional credits.

From your meeting software. If you record webinars or calls, turn on automatic transcription in the settings and a VTT file is produced alongside every recording. This is the highest-leverage version of the whole workflow: the transcript exists before you have decided you want it, for every session, without anyone remembering to do anything.

From YouTube. Uploaded videos get generated captions there too.

If you run anything on a regular schedule, turn on automatic transcription today. It costs nothing and it means every recording you make from now on is already clippable.

What Is Actually Inside the File

Worth demystifying, because the file extension makes it sound technical and it is not.

Open an SRT in a plain text editor, meaning Notepad or TextEdit rather than a word processor. On a Mac or PC you may need to right-click and choose which program opens it. What you will see is a numbered list: a timecode range, then the words spoken during it, then the next one.

That is the entire format. It is old, simple technology, and the simplicity is exactly why it works here. Every line of speech carries the moment it happened, which is precisely what you need to turn "that was a good bit" into a start point and a stop point.

The Micro-Clipping Skill

In Dubb Agent, open skills, find the general category, and switch on micro-clipping. That teaches the agent what you are asking for when you hand it a transcript, and you will get noticeably worse results without it.

Then start a conversation, attach the SRT or VTT, and describe what you want.

The Prompt

Short and specific beats elaborate here. This is the shape I use.

Mine three short clips from this webinar transcript. WHAT I WANT THEM FOR: [Social media, under 30 seconds each. Or: YouTube, two to three minutes. Say which platform and what length.] WHAT MAKES A GOOD ONE FOR ME: [A complete thought that stands alone without the surrounding context. A specific claim, a worked example, or a strong opinion. Not an introduction and not a summary.] FOR EACH CLIP RETURN: - Exact start and end timecode from the file - The transcript text for that section - One line on why it stands alone - A suggested caption or title Do not invent timecodes. Every one must appear in the file I gave you. If fewer than three segments genuinely stand alone, give me fewer and say so.

The last two lines matter more than the rest. A model asked for three clips will produce three whether or not the material contains three, and a timecode that is approximately right sends you hunting through footage, which is the manual work you were trying to avoid.

Asking why each clip stands alone is the other useful instruction. It is a quick honesty check, and a clip whose justification reads thin usually is.

Judging What Comes Back

Micro-clipping gives you timecodes and text. Read the text before you open any editor, because reading three passages takes a minute and cutting three clips does not.

What I look for, in order. Does it survive without context? The commonest failure is a segment that is excellent in the recording and meaningless alone, because it refers to something said twenty minutes earlier. Does it start cleanly? Transcript boundaries often begin mid-thought, and the fix is usually moving the start a sentence earlier. Is there a reason to keep watching after three seconds? A clip that opens on a conclusion has nowhere to go.

Adjusting a start point by a sentence is normal and quick. That is the editing this workflow leaves you, and it is the good kind.

Cutting the Clip

With a start and a stop, any editor will do. Open the source, go to the timecode, cut, export.

We use the Dubb desktop app because the next step is there too, but nothing about this workflow requires a particular tool. The value was created when the transcript told you where to cut; everything after that is mechanical.

This is worth saying plainly because it is what separates this from a clipping product. You are not buying a pipeline. You are using a transcript to make a decision, and then cutting with whatever you already have.

Removing the Silences

One pass worth running on every clip before it goes out.

Recorded conversation is full of pauses that are invisible while you are in the room and obvious in a thirty-second cut. In the desktop app this is one control: it identifies the silences and removes them when you apply it. We covered it as smart cut in the post on recording with the desktop screen recorder.

On a short clip the difference is larger than you expect, because a two-second gap in a thirty-second video is a very different proportion from the same gap in an hour. Tightening the pacing is usually the single biggest improvement available, and it takes one click.

Format and Length

The same clip serves different destinations differently, and deciding before you cut saves recutting.

Destination Shape Length Ask the Agent For
Short-form feeds Vertical Under 30 seconds One complete claim with a hook in the first line
Long-form video Landscape Two to ten minutes A section that teaches one thing end to end
Professional feeds Either One to three minutes An opinion or a specific example, not an explainer
Sales follow-up Landscape Under two minutes The passage that answers the objection you keep hearing

The last row is the one people forget. A clip does not have to go on social at all, and a two-minute answer to a recurring objection, sent to the person raising it, is often the highest-value thing in a two-hour recording.

Clipping Approaches Compared

Approach How It Selects Where Your Time Goes Best Fit
Dubb Agent from a transcript You ask for a set number against your own criteria Reading three passages, then cutting them A regular publishing habit from long recordings you already make
OpusClip and similar Generates many candidates automatically, with the editing done for you Reviewing candidates and discarding most Material with a high hit rate, and anyone who wants the edit finished for them
Descript You edit the video by editing its transcript directly Reading and cutting in one place People who want selection and editing in a single document-like tool
Watching it back yourself Your own judgment, which is the best available Two hours per two-hour recording Something that really matters, where nothing may be missed

These are fits rather than rankings. If your recordings are dense and quotable, generous candidate generation gives you more usable material than asking for three would. If they are long, conversational and mostly not clippable, which describes most webinars and trainings, precision wins because the hit rate on candidates is what decides everything.

Common Mistakes and Troubleshooting

Trying to clip without a transcript. The subtitle file is the requirement. Everything else follows from it.

Not turning on automatic transcription. One setting means every future recording is already clippable.

Opening the file in a word processor. Use a plain text editor, and right-click to choose it if your system opens something else by default.

Asking for as many clips as possible. That recreates the review problem you were avoiding. Name a number.

Trusting timecodes without checking. Tell it not to invent them, then confirm the first one lands where it claims.

Publishing on the transcript boundary. Segments often start mid-thought. Move the start a sentence earlier.

Skipping silence removal. A two-second pause in a thirty-second clip is a much bigger problem than it was in the original.

Only thinking about social. The clip that answers a common objection is worth more sent to one prospect than posted to a feed.

Best Practices I Actually Follow

Turn on automatic transcription everywhere. It is the one setting that makes this workflow free thereafter.

Ask for three. Not ten. Three you will publish beats ten you will assess.

Read the text before opening an editor. A minute of reading saves cutting things you will discard.

Tell it to return fewer if the material does not support three. An honest two is better than a padded three.

Run silence removal on everything. Biggest improvement per click available.

Keep one clip back for sales. Not everything needs to be published to be useful.

Proof: Why This Actually Works

The mechanism behind micro-clipping is that reading is faster than watching, and a transcript makes a two-hour video readable. Nothing about the model's judgment needs to beat yours; it only needs to narrow two hours to three passages you can assess in a minute.

Two patterns hold consistently across the people we help. The first concerns volume. People who ask for a small number of clips publish more over a month than people who generate many, because the second group accumulates a backlog of unreviewed candidates and stops opening it. The bottleneck was never the supply of clips.

The second concerns transcription settings. Anyone who turns on automatic transcription for recurring recordings is still clipping a quarter later, and anyone generating transcripts by hand each time tends to stop within weeks. The friction that matters is not the clipping, it is whether the transcript is already waiting.

Methodology note: these are directional observations drawn from aggregated, anonymized usage patterns across Dubb users, not a controlled study. No figures are attached to either pattern, and results vary by how quotable the source material is and how often recordings are made.

What I take from it is that this is a habit problem wearing a technology costume. The tools all work. What decides whether you publish is whether the transcript is sitting there on Monday morning.

Frequently Asked Questions

What is micro-clipping?

Micro-clipping is extracting short, self-contained segments from a long recording so each can be published on its own. A podcast, webinar, training or demo becomes several social posts or shorter videos.

It differs from editing because nothing is being assembled. The clip already exists inside the recording, and the whole job is deciding where it starts and stops.

Do I need an SRT or VTT file to clip a video with AI?

Yes. A subtitle file is the requirement, and SRT and VTT are interchangeable for this purpose. A model cannot watch your footage, but it can read every word you said along with exactly when you said it.

That turns a two-hour recording into a document, and finding the strongest passages in a document is something a model does well.

Where do I get a transcript file for my video?

Three common routes. Upload the recording to Dubb Studio, where captions are generated and the SRT can be downloaded, a paid-plan feature that does not consume extra credits. Turn on automatic transcription in your meeting software, which produces a VTT alongside every recording. Or upload to YouTube, which generates captions too.

If you record anything on a schedule, switch on automatic transcription now. Every future recording then arrives already clippable, with nobody having to remember anything.

Why not use a tool that generates clips automatically?

You can, and for some material it is the better choice. The difference is what the tool optimises for. Candidate generators aim not to miss anything, so they produce generously and leave you to sort.

That moves the work from editing to reviewing rather than removing it. If your recordings are dense and quotable, the hit rate makes that worthwhile. If they are long and conversational, working through a pile of near-misses can cost more than it saves, which is why asking for three specific clips suits webinars and trainings.

How do I write a good micro-clipping prompt?

Name the number of clips, the platform, the length, and what makes a good one for you, such as a complete thought that stands alone rather than an introduction or a summary.

Two instructions matter most: tell it not to invent timecodes, since an approximate one sends you hunting through footage, and tell it to return fewer than you asked for if the material does not genuinely contain that many. Asking why each clip stands alone is a useful honesty check.

What should I do to a clip before publishing it?

Check it survives without context, since the commonest failure is a segment that made sense in the recording because of something said twenty minutes earlier. Move the start point a sentence earlier if it opens mid-thought.

Then remove the silences. A two-second pause is barely noticeable in an hour and very noticeable in thirty seconds, so tightening the pacing is usually the single biggest improvement available.

Three Clips Is Enough

The workflow end to end: turn on automatic transcription, download the SRT or VTT after a recording, switch on the micro-clipping skill, ask for three clips with the platform and length named, read the passages, cut them, remove the silences, publish.

The part worth internalising is the number. Three clips you will actually post beats fifteen you will sort through next week, and it is the difference between a publishing habit and a folder of good intentions.

Dubb Agent handles the reading and the timecodes, and the cutting happens wherever you already work. For the strategy of how long-form and short-form feed each other, our post on the waterfall method covers the thinking, and once the clips exist, repurposing across channels covers getting them out.

If you change one thing after reading this, turn on automatic transcription for whatever you record regularly. Everything else in this article is a decision you can make in ten minutes, but only if the file is already there.

About the Author: Ruben
Ruben

CEO and Founder of Dubb and host of Connection Loop (a top 3% worldwide podcast), author of the best-selling book Click Record. Passionate about helping people succeed with video, AI, and automation. Empowers businesses to grow through innovative technology and storytelling.

Related Posts

View all posts