imstruzik.com
← Back to Notes
First draft by AI — Marcin to fact-check and edit

Note 03 · AI video editing

AI Video Editing: 7 Tasks to Hand Over Today, and 3 You Never Should

Seven editing tasks safe to hand over to AI today, three that still need a human eye, and the one-question test for telling them apart.

“AI video editing” gets pitched as one category, as if every part of editing were equally automatable. It is not. Some of it is a solved problem now, genuinely faster and just as good done by a model. Some of it only looks finished at first glance, and fixing it afterward costs more time than doing it by hand would have.

The 7 tasks worth automating

  • Auto-captions. Accurate enough for a first pass on almost any tool now; a quick proofread is all that's left.
  • Silence & filler-word removal. Cutting dead air and “um”s by hand was always tedious and mechanical — exactly what automation is for.
  • Format resizing. One master edit into vertical, square and widescreen cuts, instead of three separate manual re-edits.
  • Rough transcription & searchable notes. Turning raw footage into a searchable transcript makes finding the right ten seconds in an hour of recording instant instead of a scrub-through.
  • Noise cleanup. Background hum and room noise removal on a voiceover track, reliably, in one pass.
  • First-pass color matching. Evening out exposure and white balance across clips shot at different times is now a reasonable starting point to grade from, not a finished grade.
  • B-roll tagging. Auto-tagging a footage library by what's in frame turns a folder of unlabeled clips into something you can actually search.

Those are hours I am glad to give back. Doing all seven by hand in 2026 is a hobby, not a workflow.

Why "roughly right" is not good enough for 3 of them

  • Pacing of the cut. Where exactly a cut lands relative to a beat of music or a word of narration is a feel, not a rule, and AI's version is consistently a few frames off in a way viewers sense without being able to name.
  • UI-animation timing. Motion has to match precisely what's happening on screen — a click, a load, a state change — and AI-generated timing treats it as decorative instead of functional.
  • Story order. Which point comes first, what's held back for the end — this is structural judgment about the specific audience, not a pattern to be learned from other videos.

AI gets all three roughly right on the first pass, and “roughly right” is the exact distance between a video people finish and one they close early.

A time-cost comparison

Rough, typical numbers for a single 60–90 second SaaS product video — not a stopwatch log, but close enough to plan around:

TaskManualWith AI
Auto-captions~35 min~6 min
Silence & filler-word removal~25 min~5 min
Format resizing (3 formats)~50 min~12 min
Rough transcription & notes~25 min~2 min
Noise cleanup~15 min~4 min
First-pass color matching~35 min~10 min
B-roll tagging (per 200 clips)~2.5 hrs~15 min

Add it up across the seven and it is most of a day back per video — which is exactly why the three that stay manual are worth protecting. Giving back that much time only matters if what is left still holds attention.

How to tell which bucket a task is in

The test I use on a new project: if getting it 90% right saves time even after a manual pass to fix the last 10%, automate it. If getting it 90% right means a viewer notices the 10% that's wrong, it stays manual. Captions pass that test easily. The exact timing of a transition does not.

New notes like this go out by email first.

Get the newsletter
← Previous note · All notes · Next note →