AI Video Repurposing: Cut Recap Time to Under an Hour
AI Video Repurposing: Cut Recap Time to Under an Hour
TL;DR: I built an AI video repurposing skill in Claude Code that turns one raw session recording into a chapter outline, social clips, a full transcript, pull-quotes, and a recap draft — plus a verification pass that catches bad cuts before they ship. What used to be an evening of scrubbing footage is now something I kick off and check on later.
I run every live Secure AI Builder Bootcamp session — Kickoff, Topic + Q&A, Build Session, Demo Day — and the raw recording goes up on GroupApp fast. The written recap is supposed to follow close behind. I’m also about to record the entire AI-1001 free course as one continuous take instead of seven separate videos. Both of those mean a growing pile of raw footage with no repeatable way to turn it into something postable: finding quote-worthy moments, guessing at clip boundaries, writing a recap by hand. Doing that manually doesn’t scale past one or two sessions, so Saturday I built the AI video repurposing skill to do the first pass for me while I work on something else.
What the Skill Actually Produces
One command (/video-repurpose) against a raw recording generates:
- A timestamped chapter outline, scaffolded off the session’s own deck so the sections match what I actually planned to cover
- 3-5 candidate social clips (60-90 seconds), each with a suggested hook line in the same voice as our AI-1001 scripts
- A full timestamped transcript
- 5-8 pull-quote highlights — short, sharable, pulled verbatim
- A GroupApp recap draft, ready to edit and post
- A speaker feedback report — pacing, crutch words, dead air, with timestamps
That’s the whole first pass of content repurposing from a single input file, and none of it requires me to touch a video editor until I’m ready to actually cut clips.
Two Modes, Because Not Every Recording Is the Same
Freeform mode handles live sessions where I’m talking off a deck, not a script. Script-matched mode handles a different problem entirely: I’ve started recording the AI-1001 free course as one continuous take instead of seven separate videos, so the skill aligns the transcript against the seven pre-written scripts to find where each video actually starts and ends, and flags anything ambiguous instead of guessing. On the first real run, it caught a “surprise” ad-libbed-sounding segment that turned out to be planned content from a bonus slide deck I’d forgotten I’d written — checking the scripts folder alone would have missed it.
The Verification Step That Almost Shipped Two Bad Clips
Here’s the part that justified building this properly instead of hacking something together. The first version cut clips based on segment-level transcript timestamps, and on paper every duration matched the plan. Watching the actual output told a different story: 4 of 9 early test clips had leftover audio from the wrong topic or several seconds of trailing dead air.
So the skill now re-transcribes each candidate clip window at the word level and checks the real first and last spoken word against the planned boundaries before calling anything final. On the first day I used this for real, it caught two distinct problems in the Beta 2 Kickoff clips:
- One candidate opened 28 seconds early, picking up leftover audio from the previous topic
- Another had three internal dead-air gaps — 6.7, 9.3, and 11.7 seconds — that a simple start/end trim would have missed, needing a multi-segment cut-and-splice instead of one clean edit
Both clips looked correct by duration alone. Both would have gone out wrong if I’d trusted the first pass. This is the same “trust but verify” pattern I’ve written about before with our security testing — the automated first draft gets you 90% of the way, and the review step is what catches the 10% that would have embarrassed you.
Real Numbers From the First Two Runs
Transcription runs locally on my machine via faster-whisper at roughly 1.4x realtime — a 55-minute Kickoff recording finished transcribing in about 40 minutes, unattended. That’s the bulk of the wait, and it happens in the background while I work on something else.
The gap-detection technique scales past clip-length snippets, too. On a 27-minute AI-2001 Module 1 walkthrough, the same word-level approach found 42 separate dead-air gaps — one of them 37.9 seconds — across 26.6 minutes of raw footage, and trimmed the final cut down to 21.8 minutes. Finding and cutting 42 gaps by hand, one at a time, in a video editor is a multi-hour job done in small increments. The script built the segment list programmatically and did it as a single pass.
As with the documentation workflow automation I built earlier, the principle holds: let the AI do the repetitive first pass, keep a human doing the judgment calls.
Coaching, Not Auto-Editing, on Delivery
The speaker feedback report is the part I didn’t expect to find useful. On the Kickoff recording, it flagged “so” as a sentence-opener 102 times across 54.6 minutes — about once every 32 seconds — and showed my pace dropping to 47-87 words per minute during a live demo block versus roughly 140 wpm everywhere else, which lined up exactly with the section where I was waiting on tools to load.
I deliberately did not have it auto-cut those crutch words out of the video. “So” is a real word with real uses — “so that,” “so many,” a genuine transition — mixed in with its filler use, and an automated cutter can’t reliably tell the difference. Blind removal risks butchering a sentence that was fine. So the report gives me counts, rates, and timestamps, and I do the work of breaking the habit myself. Same logic as not letting an AI security scanner auto-remediate a finding it can’t fully verify: flag it, show your work, let a human make the call.
Key Takeaways
- A single Claude Code skill now turns a raw session recording into a chapter outline, social clips, full transcript, pull-quotes, and a recap draft — the first pass of video repurposing without opening an editor
- Word-level re-transcription as a verification step caught two clip-boundary errors that segment-level timestamps missed entirely, before they went out publicly
- Automated dead-air detection cut a 42-gap manual editing job down to one programmatic pass, and transcription itself runs at ~1.4x realtime in the background
- Speaker feedback (crutch words, pacing, dead air) is reported as coaching data, not auto-applied — the same trust-but-verify approach I’d apply to any AI output before it ships
Comments