1

00:00:00,000 --> 00:00:02,000

The complete guide to interview transcription

How to turn a recorded interview into an accurate, usable transcript, what affects quality, which route to take, and how to review the result properly.

On this page

You have a recording, a research interview, a podcast conversation, a job candidate call, an oral history session, and you need it as text. This is the complete guide to interview transcription: what makes interviews harder to transcribe than other audio, the three routes available to you, how to prepare your file, and how to review the output so the final transcript is something you can actually trust and cite. Interviews are their own category of transcription problem. Unlike a scripted lecture or a single-voice memo, an interview has at least two people, unpredictable turn-taking, interruptions, and often one participant on a worse microphone than the other. Understanding those specific challenges before you pick a method will save you more time than any single tool choice, so that is where this guide begins.

Why interviews are harder to transcribe than other audio

Most transcription difficulty comes from three factors: how many people are speaking, how cleanly they take turns, and how good the recording is. Interviews stress all three at once. Two or more voices means crosstalk, those moments where the interviewer says "right, right" over the middle of an answer, or both people start a sentence at the same time. Every transcription method, from a person with headphones to a modern speech model, loses some words in overlapping speech.

Interviews also carry dense proper nouns and jargon. A subject-matter expert will name colleagues, companies, drugs, court cases, or technical standards that don't appear in everyday speech. These are exactly the words a speech recognition model is most likely to misrender, and exactly the words you can least afford to get wrong when you quote someone. Plan from the start to verify names and terms against your notes.

Finally, interview recordings are often asymmetric. If you recorded a remote conversation, your own voice may be crisp while the other person arrives compressed through a conferencing codec, or from a laptop microphone across a room. The transcript quality will track the worst channel in the recording, not the best one. Knowing which speaker's audio is weaker tells you where to concentrate your review effort later.

Your three options: DIY, automated tools, and human services

The do-it-yourself route means playing the audio and typing it yourself, ideally with a player that supports slowdown and a keyboard shortcut to skip back a few seconds. It costs nothing but your time, and it costs a lot of that, typically several times the length of the recording, more if the audio is rough or the subject matter is unfamiliar. Its one real advantage: because you type every word, you finish knowing the material deeply, which some researchers genuinely value as a first analysis pass.

Automated transcription runs your recording through a speech recognition model and returns a draft transcript in minutes rather than hours. Modern large speech models handle clear, well-recorded interviews well, and they are dramatically cheaper than paying a person. Their weaknesses are the ones described above: overlapping speech, heavy accents the model has seen less of, distant microphones, and specialized vocabulary. An automated transcript should always be treated as a strong first draft that you review, not a finished document.

Human transcription services employ people who listen and type professionally. A good service still produces the most reliable text for genuinely difficult audio, a group discussion in a noisy café, a strong dialect, a recording with long stretches of crosstalk. The trade-offs are cost and turnaround measured in days rather than minutes, and the fact that you are sending your recording to another human being, which matters if the interview is sensitive or confidential.

The honest decision rule: if your audio is clear and your time matters, start automated and review carefully. If your audio is genuinely bad, either a human service or your own ears will need to be involved anyway, because no software can recover words that were never captured. If the interview is short and hypersensitive, DIY keeps it entirely in your hands.

Want to try it now? Upload a file to FastScribe. Your first one is free, no signup.

Preparing your recording for the best possible transcript

Transcript quality is decided mostly at recording time, but even after the fact there are useful steps. First, locate the original file at its original quality. If the interview was recorded in a video call platform or lives on a video hosting site, download the actual file to your device before doing anything else, transcription tools like FastScribe work from an uploaded file, so you need the recording itself, not a link to it. Prefer the highest-quality version you can get; a heavily re-compressed copy of a copy will transcribe worse.

Second, trim the file. Interviews often begin with several minutes of setup, mic checks, small talk, "can you hear me now", and end with goodbyes. Cutting these in any basic audio editor shortens processing, keeps you under file-size limits, and means timestamps in the transcript map onto the content you actually care about.

Third, if you have separate audio tracks per speaker (some recording setups produce these), consider whether to transcribe them individually. A single mixed file is simpler and is what most people have; separate tracks can help when one speaker constantly talks over the other, at the cost of merging the results yourself afterwards. For most one-on-one interviews, the mixed file is fine.

What you should not bother with: aggressive noise-reduction filters applied before transcription. Heavy processing can smear the consonants a speech model relies on and make results worse, not better. Light, conservative cleanup of a constant hum is reasonable; anything more, test on a short clip first.

Verbatim or clean read: decide what kind of transcript you need

Before you transcribe anything, decide what the transcript is for, because that determines the style. A verbatim transcript captures everything, false starts, "um" and "you know," repeated words, laughter, and is what qualitative researchers and anyone doing discourse or linguistic analysis usually needs, since how something was said carries meaning.

A clean-read (or intelligent verbatim) transcript removes filler and stutters, lightly smooths grammar, and keeps the speaker's meaning intact. This is what most journalists, podcasters, and hiring teams want: readable text that fairly represents what was said. If you will publish quotes, clean read is almost always the right target, with the crucial rule that you never change the substance of what a person said.

Automated tools generally land between the two: speech models tend to drop some hesitations naturally and keep others. Whichever style you need, plan to enforce it during your review pass rather than expecting the raw output to match it. Note your style decision at the top of your working document, "clean read, fillers removed, names verified", so that if you return to the transcript months later, or share it with a collaborator, everyone knows what editorial standard the text reflects.

The review pass: from raw output to trustworthy transcript

However the first draft was produced, the review pass is where a transcript becomes trustworthy, and it deserves a method. Work with the audio and the text side by side. Play the recording at slightly faster than normal speed while reading along; drop to normal or slow speed only where the text and audio disagree. This is far faster than re-listening to everything at full attention, and far safer than reading the transcript alone, where plausible-sounding errors hide easily.

Prioritize three things. First, every proper noun: names of people, organizations, places, and products, checked against your notes or a quick search. Second, negations and numbers, a missed "not" or a "fifteen" rendered as "fifty" inverts meaning while looking perfectly grammatical. Third, the passages you already know you will quote, which deserve word-perfect verification against the audio.

Add speaker labels as you go. Decide on a convention, full names on first appearance, initials after, and apply it consistently; in a two-person interview this goes quickly because turns mostly alternate. Mark genuinely unintelligible moments honestly with a placeholder like [inaudible 12:40] rather than guessing. A transcript that admits its gaps is more useful, and more honest, than one that papers over them.

Budget real time for this. For a clear one-hour interview, a focused review might take somewhere around half the recording's length; for rough audio, longer. Build that into your schedule rather than treating the automated draft as the finish line.

Output formats: TXT, SRT, VTT, and DOCX explained

Plain text (TXT) is the universal format: it opens everywhere, pastes cleanly into any editor or analysis tool, and is the right choice when the transcript feeds into coding software, a search index, or your own notes. Its limitation is that it carries no timing information and no styling.

SRT and VTT are caption formats: the transcript broken into short timed segments, each anchored to the moment in the recording where it was spoken. If your interview is part of a video you will publish, these files are what video players and platforms consume to display subtitles. They are also quietly useful for pure research: because every line carries a timestamp, an SRT or VTT file lets you jump straight from a quote back to the exact moment in the audio to verify it.

DOCX is the format for editing and sharing with people. It preserves formatting, works with tracked changes and comments, and is what you want when a transcript will be reviewed by an editor, a legal team, or an interviewee checking their quotes. FastScribe exports TXT, SRT, and VTT on all tiers; DOCX export is a Pro feature. A practical workflow is to keep a timestamped SRT or VTT as your reference copy and do your editorial work in a document file.

Key takeaways

  • Interview audio is hard because of crosstalk, asymmetric microphone quality, and dense proper nouns, know your recording's weak spots before choosing a method.
  • Choose the route by audio quality and stakes: automated tools for clear recordings you will review, human services for genuinely rough audio, DIY when total control matters more than time.
  • Always work from a downloaded, trimmed, highest-quality copy of the original file, transcription tools work on uploaded files, not links.
  • Treat any automated output as a first draft: review against the audio, verify every name and number, and mark unclear passages honestly rather than guessing.
  • Pick the export format for the job: TXT for analysis, SRT/VTT for captions and timestamped verification, DOCX for editorial workflows.
  • Know a service's retention policy before uploading, and store your own recordings and transcripts with care appropriate to their sensitivity.

Where FastScribe fits

FastScribe is one option in the automated-tools column, and it is worth being clear about where it does and does not fit. If you want the learning-by-typing benefit of DIY, or your recording is rough enough that only human ears will get it right, an automated tool is not the answer and we would rather tell you that than waste your afternoon. Where FastScribe fits well is the common middle case: a reasonably clear recorded interview that you want as text quickly, cheaply, and privately. It runs speech recognition on our own servers, and your audio is deleted the moment your transcript is ready, while the transcript text is kept for you. You can judge it on your own material before committing anything: your first file up to 50 MB and 10 minutes transcribes free with no signup, a free account raises the file limit to 100 MB, and Pro is $12 a month for files up to 2 GB and 5 hours, DOCX export, batch uploads, a priority queue, history kept forever, and unlimited files under fair use. What you get back is a transcript with TXT, SRT, and VTT exports (plus DOCX on Pro), the editorial work of reviewing, labeling speakers, and shaping quotes described in this guide remains yours.

Frequently asked questions

How do I transcribe an interview that's posted on a video platform?

Download the actual video or audio file to your device first, then upload that file. FastScribe works from uploaded files only, so a link to the video is not enough, you need the recording itself, at the best quality you can obtain.

Will the transcript tell me who is speaking?

Plan to add speaker labels yourself during the review pass. In a typical two-person interview this is quick, since turns mostly alternate; pick a labeling convention up front and apply it consistently as you check the text against the audio.

How accurate will an automated interview transcript be?

It depends heavily on your recording: clear voices on decent microphones transcribe well, while crosstalk, distant mics, and heavy jargon degrade results. No automated transcript should be published unreviewed, verify names, numbers, and any passage you intend to quote against the audio.

What happens to my recording after I upload it to FastScribe?

Your file is processed on our own servers, and the audio is deleted immediately after transcription. The transcript text is kept so you can return to it, and Pro accounts keep their history forever.

Can I try FastScribe on my interview before paying anything?

Yes. Your first file transcribes free with no signup, up to 50 MB and 10 minutes, enough to test a representative slice of your interview and judge the output on your own audio before deciding whether a free account or Pro makes sense.

Try FastScribe on your own recording

Your first file is free, no signup. Audio is deleted the moment your transcript is ready.

Drop an audio or video file here

MP3, M4A, WAV, MP4 and more. Free, no signup