1

00:00:00,000 --> 00:00:02,000

How to edit and clean up a transcript

A practical, step-by-step workflow for turning a raw machine transcript into a clean, readable document, fixing names, cutting filler, labeling speakers, and exporting it in the right format to share.

On this page

A raw transcript is a starting point, not a finished document. Even a strong speech model like the one FastScribe runs on our own servers will occasionally misspell a surname it has never heard, transcribe every 'um' and false start exactly as spoken, and hand you a wall of text with no speaker labels or paragraph breaks. None of that is a defect; it is simply what spoken language looks like when it is written down faithfully. The work of editing is deciding what your readers need and reshaping the text to serve them. This guide covers how to edit and clean up a transcript in a repeatable way: a first read-through pass, a targeted pass for names and terminology, a filler-and-false-start pass, structural work like speaker labels and paragraphing, the special rules that apply when you are editing caption files, and finally how to pick an export format (TXT, SRT, VTT, or DOCX on the Pro plan) so the cleaned version actually gets read. Budget roughly two to four minutes of editing per minute of audio for a careful clean-read edit; less if you only need it legible, more if it will be published under your name.

Start with a full read-through before you change anything

The most common editing mistake is opening the file and immediately fixing the first typo you see. Resist that. Read the whole transcript once, top to bottom, ideally with the original audio available in another tab so you can spot-check anything that looks wrong. On this pass you are not editing. You are building a map: who is speaking, where the topic shifts, which names and terms recur, and which stretches of audio the model clearly struggled with (crosstalk, distance from the microphone, background noise).

As you read, keep a scratch list of every proper noun, acronym, and specialist term you encounter, along with the different spellings the transcript uses for each one. A guest named 'Saoirse' might appear as 'Sersha' in one paragraph and 'Sasha' in another; a product called 'Kubernetes' might show up as 'Cooper Netties.' You will fix all of these in one systematic pass later, and the list is what makes that pass fast instead of endless.

Also decide, before you touch a word, what kind of transcript you are producing. A verbatim transcript keeps every hesitation and repetition because the exact wording matters, useful for research interviews or anywhere the record itself is the point. A clean-read transcript removes filler and smooths false starts so it reads like considered prose. Most transcripts shared with colleagues, published as interview write-ups, or attached to a video benefit from clean-read editing. Everything in the rest of this guide assumes you have made that call deliberately, because it changes what counts as an 'error.'

Fix names, places, and specialist terms in one systematic pass

Speech models transcribe what they hear, and an unfamiliar name is, acoustically, just a sequence of sounds. The model picks the most plausible spelling, which is often a phonetic guess. The efficient fix is find-and-replace, driven by the list you built during your read-through: search for each wrong spelling, replace it everywhere, then search for likely near-misses (the first syllable of the name, common homophones) to catch variants your list missed.

Verify spellings against a written source, not your memory, a meeting invite, an email signature, a company's own website, a paper's author list. Names are the single highest-stakes item in a transcript: a reader will forgive a stray comma, but the person you interviewed will notice their own name misspelled in the first line. The same goes for company names with unusual capitalization, drug or chemical names, and acronyms; expand an acronym on first use if the audience may not know it.

Two cautions for this pass. First, be careful with case-sensitive replacements, replacing 'mark' with 'Marc' will also mangle 'remarkable' unless your editor matches whole words only. Most text editors have a 'match whole word' option; use it. Second, when the audio is genuinely ambiguous and you cannot verify the term, mark it rather than guess: a bracketed '[unclear, 14:32]' with the timestamp is honest and lets you or a colleague resolve it later. FastScribe's SRT and VTT exports carry timestamps throughout, so exporting one of those alongside your working TXT file gives you a quick way to jump back to the exact moment in your original recording.

Want to try it now? Upload a file to FastScribe. Your first one is free, no signup.

Cut filler words without flattening the speaker's voice

Filler comes in several distinct kinds, and they deserve different treatment. Pure hesitations, 'um,' 'uh,' 'er', can almost always be deleted wholesale in a clean-read edit; nobody's meaning lives there. Discourse fillers like 'you know,' 'like,' 'sort of,' and 'I mean' are trickier: sometimes they are noise, but sometimes they hedge a claim the speaker genuinely wanted to hedge. Cut them where they add nothing; keep them where removing one would make a tentative statement sound confident.

False starts and repairs, 'We launched in, well, we actually launched in March', should usually be collapsed to the sentence the speaker landed on: 'We actually launched in March.' The test is simple: did the abandoned fragment carry any information the final sentence lacks? If not, it goes. Repetitions used for emphasis ('it was very, very slow') are a stylistic choice, not an error; keep them if they capture how the person talks.

The line you must not cross in a clean-read edit is changing meaning. Tightening 'I think we probably, um, might consider it' into 'We will do it' is not cleanup. It is misquotation. A good discipline: after each heavy edit, reread the sentence and ask whether the speaker would recognize it as something they said. If you are editing for publication and a passage needs real rewriting to make sense, paraphrase it outside the quotation marks rather than putting reworked words in the speaker's mouth.

Mechanically, find-and-replace helps here too, but use it with a lighter touch than you did for names. Searching for ' um ' and ' uh ' (with surrounding spaces) and deleting matches is safe; bulk-deleting 'like' is not, because half the matches will be legitimate uses. For anything beyond pure hesitations, edit by reading, not by pattern.

Add speaker labels and paragraph structure by hand

FastScribe gives you the words; it does not automatically label who said them. For any recording with more than one voice, plan on adding speaker labels yourself during cleanup. The practical method: play the first minute of audio to fix each voice in your ear, then move through the transcript inserting a label at every change of speaker. Use a consistent format, 'JORDAN:' or 'Interviewer:' at the start of the turn, one blank line between turns, and pick real names over 'Speaker 1' whenever you know them, because readers lose track of numbered speakers within a page.

Paragraphing matters just as much in single-speaker recordings. Spoken monologue arrives as one continuous stream, and a transcript that mirrors that is exhausting to read. Break at topic shifts, at rhetorical pivots ('But here's the thing…'), and roughly every three to five sentences even when the topic holds. If the recording had natural chapters, agenda items in a meeting, questions in an interview, promote those to headings so a reader can scan to the part they need.

This structural pass is also where you standardize small conventions: numbers (digits versus spelled out), timestamps if you are keeping any inline, how you mark non-speech events ('[laughter]', '[phone rings]'), and how you flag redactions if any part of the conversation should not be shared. Decide each convention once, write it at the top of your working file while you edit, and apply it uniformly, inconsistency is more noticeable to readers than whichever choice you make.

Editing caption files: the extra rules for SRT and VTT

If your destination is captions or subtitles rather than a document, edit the SRT or VTT export directly, both are formats FastScribe produces, and be aware that caption files have structure a plain transcript does not. Each cue consists of a timestamp range plus one or two short lines of text. The timestamps are what sync the words to the video, so the cardinal rule is: edit the text inside cues freely, but do not casually rewrite timestamps or delete cues outright unless you also intend the corresponding words to vanish from the screen.

Text edits inside cues follow the same principles as document edits, with tighter constraints. Keep lines short, around 42 characters per line and at most two lines per cue is a widely used convention, because captions are read in the moment, at the pace of the video. Cutting filler is usually even more valuable in captions than in documents: an 'um' on screen for two seconds is two seconds of reading capacity wasted. If cutting text leaves a cue nearly empty, it is fine for a cue to hold just two or three words; short cues read better than merged ones with mismatched timing.

Mind the format details. SRT cues are numbered sequentially and use comma decimal separators in timestamps ('00:01:04,500'), while VTT starts with a 'WEBVTT' header and uses periods ('00:01:04.500'). Editing in a plain-text editor is safest; word processors can silently convert quotes and dashes into characters some video players choke on. When you finish, load the file against the video and watch a few minutes at the start, middle, and end, a sync check takes two minutes and catches the errors that matter most.

Choose an export format and finish with a repeatable checklist

Match the format to the destination. TXT is the universal working format: every editor opens it, find-and-replace is fast, and it pastes cleanly into email, wikis, and notes tools. SRT and VTT are for anything video-facing, upload them alongside the video wherever it is hosted. DOCX, available on FastScribe's Pro plan, is the right choice when the transcript will circulate as a document: it preserves your headings, bold speaker labels, and spacing, and colleagues can add tracked changes and comments to it in any modern word processor.

A practical workflow that holds up across interviews, meetings, and lectures: (1) full read-through with a name-and-term list; (2) find-and-replace pass for names and terminology; (3) filler and false-start pass at your chosen verbatim or clean-read level; (4) speaker labels, paragraphs, and headings; (5) a final proof, ideally reading aloud or after a break, because you stop seeing your own errors after an hour in the same text. For caption files, add the sync check as step six.

One logistical note on timing your work. FastScribe deletes your uploaded audio from our servers immediately after transcription, while the transcript text is kept, so keep your own copy of the original recording locally. You will want it during editing to resolve unclear passages, and you cannot re-listen through FastScribe once the audio is gone. If your source is a video that lives online, download the video file to your machine first and upload that file; FastScribe works from uploaded files only.

Key takeaways

  • Read the whole transcript once before editing anything, and build a list of every misspelled name and term as you go, then fix them all in one find-and-replace pass verified against written sources.
  • Decide upfront between a verbatim transcript (every 'um' kept, for when the exact record matters) and a clean-read edit (filler cut, false starts collapsed), and never let cleanup change what the speaker meant.
  • FastScribe does not label speakers automatically; add speaker labels, paragraph breaks, and headings by hand, using one consistent convention throughout.
  • Edit SRT/VTT caption text freely but leave timestamps and cue structure intact, keep lines to roughly 42 characters, and always sync-check the file against the video.
  • Export to match the destination: TXT for working text, SRT/VTT for video captions, DOCX (Pro) for documents that colleagues will comment on, and keep a local copy of your original audio, since FastScribe deletes uploaded audio the moment your transcript is ready.

Where FastScribe fits

There are three honest ways to get from a recording to a clean transcript, and the right one depends on your audio and your stakes. Doing it entirely yourself, typing while listening, gives you total control and effectively merges transcription and editing into one pass, but it is slow going even for fast typists, and most people abandon it beyond short clips. Human transcription services are the strongest option for genuinely difficult audio, heavy crosstalk, strong accents, poor recordings, and for high-stakes uses where you want a person making judgment calls; the trade-offs are cost and turnaround measured in hours or days. Automated tools like FastScribe sit in the middle: you upload a file, our own transcription engine, running on our own servers, returns a transcript in minutes, and you spend your time on the editing pass this guide describes rather than on typing. No automated transcript arrives publication-ready, names, filler, and structure are always yours to fix, so the honest pitch is 'fast raw material,' not 'finished document.' You can judge whether that trade suits your audio for free: your first file, up to 50 MB and 10 minutes, needs no signup; a free account raises the file cap to 100 MB; and Pro, at $12/month, adds DOCX export, batch uploads, a priority queue, files up to 2 GB and 5 hours, history kept forever, and unlimited files under fair use.

Frequently asked questions

Does FastScribe label who is speaking in my transcript?

No. FastScribe transcribes the words but does not attach speaker names to them. Add labels yourself during cleanup: listen to the opening minute to learn each voice, then insert a consistent label like 'JORDAN:' at every change of speaker.

Can I get my audio back from FastScribe later if I find an unclear passage?

No. Uploaded audio is deleted from our servers immediately after transcription, though your transcript text is kept. Always retain your own local copy of the recording so you can re-listen while editing and resolve any passages you marked as unclear.

How do I clean up a transcript of a video that's hosted online?

Download the video file to your computer first, then upload that file to FastScribe. It works from uploaded files only. Your first file is free with no signup, up to 50 MB and 10 minutes; edit the resulting transcript using the passes described in this guide.

Which export format should I use for editing: TXT, DOCX, SRT, or VTT?

Edit in TXT if you want fast find-and-replace in any plain-text editor, or DOCX (available on Pro) if the document will circulate for comments and tracked changes. Use SRT or VTT only when the destination is video captions, and follow the caption-specific rules about timestamps and line length.

Should I remove every filler word from a transcript?

Only in a clean-read edit, and even then selectively. Pure hesitations like 'um' can go wholesale, but hedges such as 'sort of' sometimes carry real meaning, and for research or record-keeping a verbatim transcript should keep everything exactly as spoken.

Try FastScribe on your own recording

Your first file is free, no signup. Audio is deleted the moment your transcript is ready.

Drop an audio or video file here

MP3, M4A, WAV, MP4 and more. Free, no signup