Speaker labels: a transcript that says who said what

Painted illustration for speaker labels: three songbirds on a branch, tapes running from the branch down into one open book, under a dusk sky

Diarization separates voices; it does not identify people. What speaker labels can and cannot tell you, the four ways they fail, and why the recording, not the software, decides how good they are.

On this page

A speaker label is a tag on each stretch of a transcript marking which voice spoke it. You get them by transcribing with a tool that runs diarization, a second pass over the audio that groups the sound by voice, and then renaming the generic Speaker 1 and Speaker 2 to the real names yourself. That is the whole of speaker labels: a transcript that says who said what, from a voice-separation pass plus a minute of your own naming. The renaming is not a shortcoming to work around: no tool that only listens to a recording knows anyone's name. It hears distinct voices and numbers them in the order they first speak. Two things decide whether the result is usable, and neither is the software you pick: how the recording was made, and how the people in it took turns. Fast exchanges with short turns are where every automated labeler is weakest, including this one.

What speaker labels are, and what they are not

Transcription answers what was said. Diarization answers who said it, and it is a separate job done by a separate model. The two get joined at the end, which is why the feature so often sits behind its own switch or its own line on a bill: it is a second pass over the same audio, not a setting on the first one.

The important limit is the one the marketing language hides. Diarization separates voices. It does not identify people. It builds a rough acoustic fingerprint for each voice it hears, clusters the recording by fingerprint, and hands you Speaker 1, Speaker 2, Speaker 3. It has no idea that Speaker 2 is your head of sales. You supply that, once, from the recording. Anyone promising a transcript that names people from the audio alone is either describing voice enrolment against profiles you registered in advance, which is a different feature, or overselling.

That distinction matters most where the stakes are highest. If attribution has to hold up to challenge, in a deposition or anything filed with a court, an automated label is not evidence of who spoke. Transcribing recorded depositions you own covers where that line sits.

How to get speaker labels on a FastScribe transcript

The gate depends on your plan, and it is worth knowing before you upload the file that matters. Your first file needs no account and is labelled automatically, with no switch to find. It can run up to 10 minutes and 50 MB, and you get one of these per rolling seven days. You can put real names on that transcript without signing in for anything. Keep the link, though: with no account there is no history to find it in again.

On a free account there is a switch on the uploader marked Speaker labels, and you get one labelled file a day. The count resets at midnight UTC, and a file that fails to process does not spend it. The switch defaults to on, so if only one of today's recordings has two voices in it, turn the switch off for the others or the credit goes to whichever you upload first. Free accounts allow 5 files a day at up to 30 minutes and 100 MB each, and keep transcripts for 7 days.

Pro, at $12 a month, labels every file, with no daily cap, files up to 5 hours and 2 GB, batch uploads, DOCX export and priority in the queue.

Renaming costs nothing on any plan. Open the transcript, click a speaker heading, and type a real name, up to 60 characters. On an account you sign in first; on the no-account file, whoever holds the link can do it, because that transcript has no owner to sign in as. The names then ride into every export: TXT, SRT, VTT, and DOCX on Pro. Rename once and it is done everywhere, with no find-and-replace pass across downloaded files.

One note on what is kept. The audio file is deleted once the transcript is ready. The transcript text itself stays in your history so you can come back to it, which is what makes renaming and re-exporting possible at all.

Want to try it now? Upload a file to FastScribe. One file a week is free, no signup.

Where speaker labels go wrong

Four failure shapes cover almost everything, and they are all shapes of the recording rather than bugs in the software.

Fast exchanges with short turns. This is the structural weakness of automated labeling, and the honest headline. A quick "right", "yeah", "mm-hm" dropped into the middle of someone else's sentence is under a second of audio with no clean edges, and it will often be absorbed into the surrounding speaker rather than attributed to the person who actually said it. FastScribe smooths runs under about six tenths of a second that sit between two runs of the same other speaker, on the reasoning that a boundary landing mid-word is more likely than a genuine interjection at that scale. It only does this when the short run butts straight against the speech on both sides. A brief remark with a clear pause around it is deliberate speech and keeps its own label. So the trade is narrower than it sounds: you get a much cleaner transcript, and you lose the back-channel that lands inside somebody else's words.

Crosstalk. When two people speak at once, there is one waveform carrying two voices. The labeler has to pick one. Expect either a flip mid-sentence or both lines given to whoever was louder.

Similar voices in the same room on one microphone. Two people with close pitch and cadence, recorded from a laptop in the middle of a table, can merge into a single label.

Very brief participants. A voice that speaks once and holds less than about three seconds gets absorbed into its nearest neighbour. That rule exists because outro music, a door slam, laughter and echo otherwise get minted as a confident Speaker 4. Coming back is what saves you: two or more separate turns keep their own label on about half a second of speech in total, because an artifact is one burst at one point in the file while a real quiet participant speaks again later. So the colleague who chipped in six times keeps a label, and the one who said a single short sentence in an hour probably does not.

Group recordings stack all four at once, which is why focus group transcription needs more review than a two-person call, whatever tool you use.

How to record so speaker labels come out right

Separation happens at the microphone, not in the model. Everything useful you can do happens before you press record, and most of it costs nothing.

Give each person their own microphone if you can, even into one mixed recording: close, clean voices separate far better than distant ones. Some call platforms can record each participant on a separate track, which is cleaner still, but FastScribe transcribes one file at a time and will not interleave several tracks into a single conversation, so you would be lining them up by timestamp yourself. In a room, put the recorder between the speakers rather than next to one of them, and get it off the hard table surface. Avoid speakerphone, which mixes everyone into one compressed channel.

Then two behavioural things that cost nothing. Ask everyone to say their name in the first minute, which gives you a reliable key for mapping numbers to people afterwards. And ask people to leave a beat before replying. A half-second gap between turns is the difference between clean edges and constant guessing. It feels stilted for about three minutes and then nobody notices.

How to prepare audio so it transcribes accurately covers the rest of the recording side, most of which helps the words and the labels at the same time.

Fixing speaker labels after the transcript arrives

Build the map before you touch anything. Play the first minute or two, note that Speaker 1 is the interviewer and Speaker 2 is the subject, and only then rename. People who rename as they read tend to get halfway through, hit a passage where the labels swapped, and start second-guessing the whole document.

Once the names are in, spot-check the places you already know are hard: the fastest exchange in the recording, any moment where two people laughed at once, and the first thirty seconds, where a labeler has heard the least and is least sure. Fix those by hand. Do not proofread the whole file for label errors unless you are going to publish it; a searchable internal record does not need that pass. How to edit and clean up a transcript covers the rest of the cleanup, including how much of it is worth doing for a given purpose.

If you transcribed a file without labels and want them, upload it again with labels on. On a free account that spends the day's credit, so it is worth deciding before the first upload rather than after.

When speaker labels are not the answer

Say plainly where this does not fit, because it saves you an afternoon. If you need captions or notes while a meeting is happening, this is not the tool. Nothing here runs while you talk. You upload a finished recording and get a transcript back.

If you need the system to recognise a specific known person across many recordings, matching a voice against a profile you registered, that is voice enrolment and this product does not do it. Labels are per file, and Speaker 1 in Monday's recording has no relationship to Speaker 1 in Tuesday's.

If attribution has to be defensible, use a person. And if your recording is twelve people around a table captured on one phone, no automated labeler will rescue it. Budget for a human transcriber, or re-record with better microphone placement if the conversation can be repeated. The complete guide to interview transcription covers choosing between the routes for a recording you cannot repeat.

Key takeaways

  • Diarization separates voices and numbers them in speaking order; it never knows a name. You supply the names once and every export carries them.
  • Fast exchanges with short turns are the structural weak case for every automated labeler, so spot-check the quickest exchange in the recording.
  • Separation happens at the microphone, not in the model: close voices, a recorder placed between speakers, and a beat left between turns do more than any setting.
  • Ask everyone to say their name in the first minute of the recording; that gives you the key for mapping numbers to people afterwards.
  • Build the speaker map from the first minute or two before renaming anything, then fix the known-hard spots rather than proofreading the whole file.
  • A voice that speaks once for under about three seconds is absorbed on purpose; a quiet participant who speaks again later keeps their own label.

Where FastScribe fits

FastScribe labels speakers with a diarization pass that runs on our own hardware alongside transcription; no third-party AI service receives your audio, and the audio itself is deleted the moment the transcript is ready. Your first file needs no account and is labelled automatically, up to 10 minutes and 50 MB, one per rolling week. A free account gives 5 files a day at up to 30 minutes and 100 MB each, one of them labelled, chosen with the switch on the uploader, with transcripts kept for 7 days. Pro is $12 a month for files up to 5 hours and 2 GB, labels on every file, batch uploads, DOCX export and priority in the queue. Renaming speakers costs nothing on any plan, including the no-account file, and the names carry into TXT, SRT and VTT exports, and DOCX on Pro. None of this identifies anyone by name from the audio, and fast exchanges with short turns are where the labels are weakest, here as everywhere. The transcript text stays in your history so renaming and re-exporting stay possible, which also means it sits on our servers and can be read there; material you would not want a second person to read does not belong on any hosted service, this one included.

Frequently asked questions

Do speaker labels tell me who is speaking by name?

No. The system separates the voices it hears and numbers them Speaker 1, Speaker 2 and so on, in the order they first speak. It has no way of knowing a name from audio. You rename them yourself on the transcript page, free on every plan and on the file you run without an account, and the names carry into every export.

Are speaker labels free?

Partly, and the shape matters. Your first file with no account is labelled automatically with no switch to set. A free account gets one labelled file a day, chosen with a switch on the uploader, resetting at midnight UTC. Pro at $12 a month labels every file. Renaming a speaker is free everywhere, including the file you run without an account.

Do speaker names appear in downloaded files?

Yes. Once renamed, the names appear in TXT, SRT and VTT exports, and in DOCX on Pro. Rename once and every format you download afterwards carries it.

Can I tell it how many people are in the recording?

No. It works out the number of voices itself, and there is no field for expected speaker count. Recordings with many voices and short turns are the hardest case, so if you know a file is a six-person discussion, plan on a review pass rather than a setting.

Why did two people end up with one speaker label?

Almost always similar voices recorded on a single distant microphone. The labeler separates voices by their sound, so two people with close pitch and cadence, picked up by the same laptop from across a table, can cluster as one. The reverse also happens: one person who moved around the room can split into two labels. Speaker labels are strongest on a two-person conversation with a microphone each, and weakest on fast exchanges with short turns.

Try FastScribe on your own recording

One free transcription a week, no signup. Audio is deleted the moment your transcript is ready.

audio or video · one free transcription a week, no signup