1

00:00:00,000 --> 00:00:02,000

How to prepare audio so it transcribes accurately

The quality of a transcript is mostly decided before you ever upload the file, here is how to record so the words come back right.

On this page

Automatic speech recognition has improved dramatically, but it still obeys an old rule of audio engineering: garbage in, garbage out. A muddy recording made across a reverberant conference room will produce a muddled transcript no matter which engine processes it, while a clean, close-miked recording of a single steady voice will sail through almost anything. The encouraging part is that most of what separates those two outcomes costs nothing. It is a matter of where you sit, how you set up, and a few habits you adopt before pressing record. This guide explains how to prepare audio so it transcribes accurately, working from the room outward: acoustics, microphone placement, recording levels, speaker habits, technical settings, and a short pre-flight test that catches problems while they are still fixable. Everything here applies whether you eventually transcribe the file yourself, send it to an automated service, or hand it to a professional.

Start with the room, not the microphone

Speech recognition systems are trained overwhelmingly on clear, direct speech, so the single biggest enemy of a good transcript is not a cheap microphone. It is the room. Hard parallel surfaces (bare walls, glass, tile, long conference tables) bounce sound back at the microphone a few milliseconds late, smearing consonants together. To a human listener this reads as 'echoey'; to a recognition model it reads as ambiguity between similar words, and ambiguity becomes errors on the page.

Before you consider any gear upgrade, choose the softest, smallest reasonable space available. A furnished office with carpet, curtains, and a couch will outperform a glass-walled meeting room every time. Recording at home, a closet full of clothes is a running joke among podcasters precisely because it works. If you are stuck with a hard room, get closer to the microphone (more on that below), proximity is the cheapest form of acoustic treatment there is.

Also hunt down steady background noise before it embeds itself in your file. Air conditioning, refrigerators, fans, an open window over traffic, continuous broadband noise raises the floor under the speech and forces the model to guess at quieter syllables. Turn off what you can, close what you can, and move away from what you cannot. Thirty seconds of walking around the space and simply listening will reveal more than any spec sheet.

Get the microphone close, and keep it there

Distance from mouth to microphone is the most powerful variable you control. Every doubling of that distance drops the direct level of the voice by roughly six decibels while the room noise and reverberation stay constant, so the ratio of speech to everything-else collapses quickly. A modest microphone fifteen to thirty centimeters from the speaker will beat an expensive one sitting two meters away on a table, reliably and by a wide margin.

For a single speaker, a headset or lavalier microphone is ideal because the distance never changes as the person turns their head or leans back. A USB desk microphone works well too, provided the speaker stays anchored in front of it. Laptop built-in microphones are a last resort: they sit near the fan, point at the keyboard, and often apply aggressive processing designed for calls rather than recording.

For a group around a table, resist the urge to put one device in the middle and hope. If a dedicated boundary or conference microphone is not available, place the recorder as close to the centre of conversation as physically possible, ask people to lean in when speaking, and accept that whoever sits farthest away will need to project. When a meeting really matters, recording a second backup device closer to the quietest participant is cheap insurance.

Want to try it now? Upload a file to FastScribe. Your first one is free, no signup.

Set levels once, then leave them alone

Recording level, gain, has a sweet spot, and both failure modes hurt a transcript. Set it too hot and loud passages clip: the waveform slams into the digital ceiling and the peaks are flattened into distortion that no downstream processing can undo. Set it too low and you will later amplify the file, dragging the noise floor up along with the voice, which is exactly the ratio you spent all that room effort improving.

A practical target on most recorders and interfaces is speech peaking around minus twelve to minus six decibels on the meter, with the loudest laughter or emphasis still leaving headroom below zero. Do a quick soundcheck at conversational volume, not a shy 'testing, testing', and set gain against how people will actually talk. Digital recordings have enormous dynamic range, so when in doubt, err slightly quiet rather than risk clipping.

Be wary of automatic gain control on phones and consumer apps. It is built to keep a voice call intelligible, and it does that by pumping the level up during silences, which drags hiss and room rumble into every pause. If your recording app lets you switch to a manual or 'music'-style mode, that fixed gain will usually give a cleaner file for transcription purposes.

Coach the speakers before you press record

Overlapping speech is where automatic transcription suffers most. When two people talk at once, their words interleave in a single waveform and any system, human or machine, must reconstruct two sentences from one signal. A ten-second word from the person running the session ('try to let each other finish; the recording is going to be transcribed') measurably improves the output, and it also produces a better conversation.

Names, acronyms, product codes, and technical jargon are the other predictable trouble spots, because a recognition model falls back on statistically likely words when it hears something rare. Ask speakers to say unusual terms clearly the first time, and consider spelling out critical ones ('that's Kowalczyk, K-O-W-A-L-C-Z-Y-K') on the recording itself. You will thank yourself when correcting the draft, since a term transcribed wrong once tends to be transcribed wrong the same way throughout.

Pace matters less than people fear. Modern systems handle brisk natural speech well, but trailing off matters a lot. Sentences that die into a mumble, asides delivered while turning away from the microphone, and words spoken through a yawn or a sip of coffee all vanish first. Encourage speakers to finish their sentences at the volume they started them.

Choose sensible technical settings

The good news is that speech does not demand audiophile settings. A sample rate of 44.1 or 48 kHz with 16-bit depth captures everything a transcription engine can use. Mono is perfectly fine for a single microphone; there is no benefit to a stereo file that contains the same voice in both channels, and it doubles the file size for nothing.

Format matters less than generation loss. Recording directly to a lossless format such as WAV or FLAC is safest; a high-bitrate compressed file (for instance a good-quality MP3 or M4A straight from the recorder) also transcribes well. What genuinely hurts is re-encoding: taking an already-compressed file and converting it again, or extracting audio through a chain of apps that each transcode it. Every lossy pass discards detail, and consonants, the part of speech that distinguishes words, go first.

Keep the original file untouched until the transcript is done and checked. If you trim, denoise, or convert, work on a copy. And if your source is a video, you rarely need to extract the audio at all, most transcription workflows accept common video files directly, and pulling the audio out yourself only adds an opportunity to degrade it.

Run a thirty-second test before the real thing

Almost every disastrous recording could have been caught by a short rehearsal. Record thirty seconds of the actual setup, same room, same seats, same microphone, people speaking at natural volume, then listen back on headphones, not the laptop speaker. Headphones expose hum, hiss, and echo that small speakers politely hide.

Better still, run that test clip through the same transcription process you plan to use for the real session. If the draft comes back clean, you have validated the entire chain end to end. If names are mangled or a distant participant disappears from the text, you have found out while there is still time to move a chair, raise the gain, or shut a door.

Make the check a habit with a fixed mental list: room quiet, microphone close, levels peaking in the healthy zone, storage space sufficient, device plugged in or fully charged, and 'do not disturb' enabled on any phone doing the recording. It takes two minutes and it is the highest-leverage two minutes in this whole guide.

Salvaging recordings you have already made

Sometimes the recording already exists and preparation is off the table. There is still a sensible order of operations. First, listen and diagnose: is the problem low volume, constant noise, echo, or overlapping voices? Each has a different remedy, and applying the wrong one makes things worse.

Low overall volume is the easiest fix, normalize or amplify the file in any free audio editor so speech sits at a healthy level. Steady background noise responds reasonably well to gentle noise reduction, but use a light touch: aggressive settings produce a warbling, underwater artifact that confuses recognition models more than the original hiss did. Reverberation and crosstalk, unfortunately, are largely baked in; expect to spend more time correcting those transcripts by hand.

Whatever cleanup you do, export the processed copy in a lossless format and try transcribing both versions. Occasionally the untouched original outperforms the 'improved' one, because the processing removed cues along with the noise. Comparing a paragraph or two of each output takes minutes and tells you definitively which file to commit to.

Key takeaways

  • The room and the microphone distance decide more of your transcript quality than any software choice made afterward.
  • Keep the microphone within about thirty centimeters of the speaker; proximity beats equipment price almost every time.
  • Set gain so speech peaks around minus twelve to minus six decibels. Clipping is unrecoverable; slightly quiet is fine.
  • Brief the speakers: one voice at a time, finish sentences fully, and say rare names or jargon clearly the first time.
  • Record lossless or high-bitrate, avoid re-encoding, and always keep the untouched original file.
  • Test thirty seconds of the real setup, listen on headphones, and transcribe the test clip before the session that counts.

Where FastScribe fits

Once the file is recorded, you have three realistic routes to a transcript. The do-it-yourself route, running an open-source speech model on your own computer, costs nothing and keeps audio entirely on your machine, but it requires comfort with software setup and reasonably capable hardware, and long files can take a while. Human transcription services sit at the other end: they remain the best choice for genuinely difficult audio (heavy crosstalk, thick accents over bad phone lines, exacting formatting requirements), at the price of higher cost and slower turnaround. Automated services occupy the middle, and FastScribe is one of them: you upload a file, our own servers run it through our own transcription engine, and you get a transcript you can export as TXT, or as SRT/VTT caption files (DOCX export is part of Pro). Your first file, up to 50 MB or 10 minutes, needs no signup; a free account raises the ceiling to 100 MB; and Pro, at $12/month, allows files up to 2 GB or 5 hours, with batch uploads, a priority queue, history kept forever, and unlimited files under fair use. Audio is deleted immediately after transcription; the text is kept. Note that it works on uploaded files only. There is no pasting a link or connecting an account, and it produces a transcript, not summaries or polished documents. If your recordings are clean because you followed this guide, an automated service like ours is usually the fastest, cheapest fit; if they are rough, weigh the DIY and human options honestly.

Frequently asked questions

Can I transcribe the audio from an online video?

Yes, but you need the actual file. Download the video to your computer first, then upload that file to whichever transcription tool you use, most accept common video formats directly, so extracting the audio separately is usually unnecessary.

Does background music ruin a transcription?

Music mixed under speech is one of the harder cases for recognition systems, because it overlaps the same frequencies as the voice. If you control the recording, keep music out entirely and add it back later in editing.

Should each speaker have their own microphone?

When practical, yes. Separate close microphones keep every voice loud and clear regardless of seating, and if your recorder captures separate tracks you can transcribe each one individually, which sidesteps most problems caused by overlapping speech.

What file format should I upload for the best result?

A lossless file such as WAV or FLAC is ideal, but a first-generation high-bitrate MP3 or M4A performs nearly as well. The thing to avoid is re-encoding a compressed file a second time, which degrades the consonants that distinguish words.

Is it worth denoising a recording before transcribing it?

Only gently, and only for steady background noise like hiss or hum. Aggressive noise reduction creates artifacts that confuse recognition models, so process a copy, compare transcripts of both versions, and keep whichever reads cleaner.

Try FastScribe on your own recording

Your first file is free, no signup. Audio is deleted the moment your transcript is ready.

Drop an audio or video file here

MP3, M4A, WAV, MP4 and more. Free, no signup