From Two-Hour Interview to Paper Edit: Running a Transcript Through a Chat Assistant | Cutroom AI
Home Your AI Assistant From Two-Hour Interview to Paper Edit: Running a Transcript Through a Chat Assistant

From Two-Hour Interview to Paper Edit: Running a Transcript Through a Chat Assistant

A practical workflow for turning a long unscripted interview into a story-ordered paper edit with a chat assistant, without losing timecode accuracy.

By Hana Okafor, a post-production supervisor for unscripted TV · Published 27 May 2026 · 9 min read · Reviewed against our editorial standards

ADVERTISEMENT

The paper edit is where an unscripted show is actually written. Nobody sees it, but the difference between an episode that breathes and one that flails usually traces back to whether someone sat with the transcript and found the spine of the story before an editor touched the timeline. On the shows I supervise, that job has historically eaten a full day per hour of interview. A chat assistant will not do it for you, but it can take the mechanical drudgery out of the first two passes so you spend your day on judgment instead of scrolling.

Here is the workflow I actually use, including the parts that go wrong.

Start with a transcript that carries timecode

The single most important decision happens before any AI touches the words. Your transcript has to carry timecode, and it has to carry it in a form the assistant will not silently mangle. I transcribe with a tool that exports word-level or segment-level timecode — in 2026 that is usually the transcription built into Premiere or Resolve, or a standalone pass through something like Descript or a Whisper-based service. What I want out is a plain text or SRT/VTT file where every paragraph or line is stamped, like this:

[00:14:22] So the first time I realized the business was failing,
I was standing in the walk-in freezer counting boxes.

[00:14:31] And my brother calls, and I just... I couldn't tell him.

Do not paste in a transcript with no timecode and expect the assistant to invent one. It will happily produce timecodes that look plausible and are completely fabricated. That is the failure mode that ends with an editor pulling their hair out at 9pm because the selects don't exist where the paper edit says they do. Timecode is ground truth; it has to come from your transcription tool, not the language model.

Pass one: an honest content map

My first prompt to the assistant is deliberately not "make me a paper edit." It is a summarizing pass that forces the model to prove it read the whole thing. I paste the timecoded transcript (a two-hour interview is roughly 18,000 to 22,000 words, which fits comfortably in the context window of Claude or ChatGPT in 2026) and ask for a topic map:

You are helping me, a post-production supervisor, build a paper edit
from a raw interview transcript. The transcript below has timecodes
in [HH:MM:SS] format at the start of each segment.

First pass only: give me a chronological content map. For each
distinct topic the subject covers, output:
- the starting timecode
- a 1-line description of what they talk about
- an emotional temperature (flat / warm / heated / emotional)

Do not paraphrase quotes yet. Do not reorder anything. Use ONLY
timecodes that appear in the transcript. If you are unsure where a
topic starts, say so rather than guessing.

TRANSCRIPT:
[paste]

That "use only timecodes that appear in the transcript" line matters, and I still spot-check it. The emotional temperature column sounds soft but it is the most useful thing on the page when I am hunting for act-outs and buttons later.

Pass two: pull the selects for a specific story

Now I know what is in the tape. The second pass is where I bring the story I actually want to tell. Unscripted interviews are sprawling; your episode has a thesis. I tell the assistant the thesis and ask it to pull candidate soundbites in service of it.

The episode's story is: a family restaurant owner nearly loses the
business during a bad year, and rebuilds it by bringing his
estranged brother back in.

Using the content map and transcript, pull the strongest soundbites
that build THIS story. Group them into a rough three-act structure
(setup / low point / resolution). For each selected bite give me:
- in and out timecode
- the verbatim quote
- one line on why it earns its place

Keep quotes verbatim. Flag any bite where the subject rambles and
might need a mid-sentence cut, but do not rewrite their words.

The verbatim instruction is non-negotiable and you must verify it. Language models paraphrase by reflex — they will smooth "I couldn't, I just, I didn't tell him" into "I couldn't bring myself to tell him." That cleaned-up version is a lie about what is on your tape. When you build the actual paper edit, every word has to be a word the subject said, in an order they said it (or a defensible frankenbite you can hear working). I read every returned quote against the transcript. This is faster than writing selects from scratch, but it is not unsupervised.

Pass three: the paper edit itself

With vetted selects in hand, I ask for the paper edit as an ordered document an editor can cut from. I want it in a form that pastes cleanly into a marker list or a doc the editor keeps open beside the timeline.

I explicitly ask it not to invent transitions or VO that doesn't exist. If a gap needs a bridge, I want it flagged as [BRIDGE NEEDED], not filled with imaginary narration.

Where it saves time, and where it will burn you

Honest accounting, because the tool is genuinely useful and genuinely dangerous in specific ways.

It saves real time on: the content map, which is pure grunt work; finding every place the subject touches a given theme across two hours, which humans miss when they get tired; and producing a clean, consistently formatted document instead of the ransom-note selects most of us type at midnight.

It will burn you on: timecode drift, if your source transcript is sloppy; fabricated or smoothed quotes, if you don't verify; and false confidence about tone. The model reads words, not performance. It cannot hear the pause before the subject's voice cracks, and that pause is often the whole shot. It also has no idea what you covered in B-roll or what the director promised the network, so its act breaks are suggestions, not decisions.

There is also a real privacy consideration. Interview transcripts frequently contain private information about people who did not consent to being fed into a third-party model. Check your production's data policy and the platform's data-retention terms before pasting anything sensitive. Many shops now run a self-hosted or enterprise-tier model with no-training guarantees specifically for this reason. If you are on a show with an NDA or medical, legal, or minors' content, treat that as a hard stop until legal signs off.

What the assistant is actually for

The paper edit is a creative act, and the assistant does not do the creative part. It does the reading, the finding, and the formatting — the three things that make the creative part exhausting. When I hand an editor a story-ordered document with verified timecode on day one instead of day three, they start cutting picture two days earlier, and the story is better because I spent my day choosing bites instead of hunting for them. That is the whole value. Keep the machine on the grunt work, keep your ears on the tape, and never let a paraphrased quote reach the timeline.

transcriptspaper-editunscriptedworkflow

Put this into practice

Compare a flat monthly chat subscription against the equivalent API usage and find the break-even point where one overtakes the other.

Open the Subscription vs API Cost Comparison →

A note on shelf life. AI products change fast. This guide deliberately focuses on the parts that stay true — how to judge a tool, what the trade-offs are — rather than ranking products that will have changed by the time you read it. Prices and feature claims should always be checked against the provider before you rely on them.