Research

What is speaker diarisation? How transcripts separate different voices

Dylan de Heer, Co-Founder & CPO

Dylan de Heer

What is speaker diarisation? How transcripts separate different voices

Speaker diarisation divides an audio recording into segments associated with different speakers. It answers a practical question: which voice was speaking at each point? Combined with transcription, it turns a block of recognised words into a conversation with speaker labels and timestamps.

You will also see the American spelling, speaker diarization. Both refer to the same task. The labels might initially be “Speaker 1” and “Speaker 2”; assigning a real person's name is a separate step that needs supporting information.

That distinction matters when an interview becomes research evidence or a meeting becomes an action list. Correct words attached to the wrong person can still produce a misleading record. NVIDIA's speaker diarisation overview explains the separation between recognising words and determining the speaker timeline.

A simple example: the words are right, but who said them?

Consider this fictional exchange:

  • 00:04 — Speaker 1: “Could you send the revised brief on Friday?”

  • 00:09 — Speaker 2: “I can send a draft, but the final version needs review.”

  • 00:16 — Speaker 1: “That works. Let's review it together.”

A transcript without labels preserves the sentences, but a reader has to infer the turns. A diarised transcript makes the separation explicit. After checking the recording and the participants, a reviewer might identify Speaker 1 as Alex and Speaker 2 as Robin.

Now imagine the second sentence is incorrectly assigned to Alex. The words remain unchanged, yet a summary could give Alex responsibility for sending the draft. There is another separate risk: changing “a draft” into “the final version”. Speaker attribution and wording both need review before extracting a commitment.

This is why diarisation is useful without being a guarantee of an accurate meeting record. It adds structure that you can inspect. It does not remove the need to inspect consequential statements.

Diarisation, transcription and recognition do different jobs

Several speech technologies are often grouped under “AI transcription”. Understanding their boundaries helps you choose the right feature and diagnose the right problem.

Transcription, or automatic speech recognition, produces words. Its question is what was said. A transcript can be readable even when it has no speaker labels.

Speaker diarisation groups speech by speaker over time. Its question is which parts appear to belong to the same voice. The output can use anonymous labels. A system does not need to know a person's name to distinguish their turns from someone else's.

Speaker identification or recognition attempts to connect a voice to an identity. Depending on the product, that may involve previously enrolled voices, stored information or manual confirmation. A familiar name appearing beside a passage is an additional claim beyond anonymous diarisation.

Speech separation attempts to recover distinct audio signals from a mixture. Speaker labels in a transcript do not automatically give you a clean audio track for each person. If two people talk at the same time, a labelled passage and an isolated voice recording are different outputs.

NVIDIA describes the diarisation task in terms of speaker-labelled audio segments. AssemblyAI's technical explainer also distinguishes diarisation from identifying a known speaker and separating mixed speech. Product interfaces may combine these tasks, so check what the feature actually returns.

How speaker diarisation works

A common pipeline performs several stages. The exact architecture varies; this is a conceptual explanation rather than a claim that every product uses identical steps.

Find the parts of the recording that contain speech

Voice activity detection identifies intervals likely to contain speech. This helps later stages concentrate on voices rather than treating every silence or background sound as a speaker turn.

The distinction is useful when reviewing an output. A missing utterance is not necessarily a naming problem. It may have been missed before any speaker label was assigned, or the words may have been lost during transcription.

Represent the characteristics of each voice

The system analyses speech segments and builds numerical representations of voice characteristics, commonly called speaker embeddings. These representations support comparisons between segments; they are not the transcript text and are not, by themselves, a person's name.

Group segments that appear to come from the same speaker

A clustering stage can group similar voice representations. The system then assigns a speaker label to the relevant time intervals. Some systems estimate how many speakers are present; others can use a supplied count.

The output is a timeline. When combined with word timestamps from transcription, that timeline can support a readable sequence of labelled turns. The words and speaker boundaries still come from different estimates and can disagree around a transition.

Recognise that other architectures exist

Not all systems expose these stages separately. NVIDIA documents both cascaded pipelines, which include voice activity detection, embeddings and clustering, and end-to-end systems that predict speaker activity using a single model. See its diarisation system types.

For someone reviewing meeting notes, the practical questions are usually simpler than the architecture: can you inspect the source, correct a mistake and tell which identity assignments are confirmed?

Where speaker labels can mislead

Do not judge the whole recording from its first few neat turns. Review the places where the result matters and where the structure looks inconsistent.

One person appears under several labels. Listen to clear passages for each label before merging anything. Several labels may represent one voice, but they may also represent people who sound similar. A tidy transcript is not sufficient reason to collapse them.

Different people share one label. Check a passage where you know a handover occurred. If two clearly different voices have been grouped together, renaming the shared label will not solve the underlying separation problem.

A brief reply is attached to the previous speaker. Small utterances such as “yes”, “no” or “I can” can matter disproportionately. Replay the transition and the surrounding sentence before using the reply as evidence of agreement or ownership.

Two people overlap. Flag the passage for review rather than forcing a confident sequence that the audio does not support. Diarisation and speech separation are different tasks, and the available recording may not allow a reliable attribution.

A plausible name hides an uncertain match. Confirm how that name was assigned. The participant list, first person to speak or person who talks most does not establish the identity of every labelled turn.

These are review patterns, not measured error rates for a particular product. Recording conditions and system choices affect performance. Do not transfer a vendor's result on one dataset to your own interviews without checking comparable material.

How to check a diarised transcript before using it


Cream cloth headphones beside a separate apricot cloth magnifying glass

Start with a clear passage from each apparent speaker. Listen to enough surrounding audio to distinguish a voice from a one-word response. Confirm names only where the recording or reliable participant information supports the assignment.

Then inspect the parts you plan to reuse. For a research interview, that may mean a quotation and the question preceding it. For a project meeting, prioritise decisions, commitments, objections, dates and amounts. A correct label at the beginning does not establish that every later turn is correct.

Use the following sequence:

  1. Check the words. Does the transcript preserve the meaning, including qualifiers such as “if”, “possibly”, “draft” and “not yet”?

  2. Check the turn. Does the quoted span belong to one speaker, or does it cross a change of voice?

  3. Check the identity. Is the displayed name confirmed, manually assigned or still an uncertain match?

  4. Check the context. Was the statement a proposal, question, refusal or commitment? Read and listen around it.

  5. Check the downstream wording. Does the summary or action list retain the same person, scope and uncertainty?

If an attribution remains unclear, keep an anonymous label or mark the passage as uncertain. Where necessary, ask the participants to clarify the intended responsibility rather than using the software's label to settle it.

For a fuller interview workflow, see how to transcribe a research interview. That process includes checking the wording as well as who spoke it.

Using speaker-labelled source material in Weeve

Weeve performs speaker diarisation on the Mac as part of its recording processing. Its current implementation combines the speaker timeline with word timings to build labelled transcript segments. Speaker assignment and people information are separate parts of the workflow; a readable name should still be checked when the attribution matters.

The speaker identification feature overview describes the product context. For the underlying product change, the FluidAudio speaker-detection update explains why this stage exists in Weeve's recording workflow.


Weeve Chat with a transcript source panel showing named speakers and timestamps in demonstration content

Genuine Weeve interface with demonstration content. The source panel shows how an answer can be checked against speaker-labelled transcript passages.

When using Chat with recordings, open the relevant source material before relying on a statement about who agreed to do something. A useful question is: “Which passages support this action, who is labelled as speaking, and what remains uncertain?” Treat the answer as a route back to the evidence, then check the recording where available.

Keep processing choices separate too. Local transcription and diarisation do not mean that every optional feature in every configuration stays on the device. Choosing an external AI provider for a later chat changes the processing boundary for the material sent to that provider. Check the selected mode before using sensitive content.

Choosing a tool: questions more useful than a headline accuracy claim

Evaluate the output you need. For an interview, you may want a transcript with timestamps and editable speaker labels. For an audio editor, separate voice tracks may be essential. Those are different requirements even if both products advertise speaker features.

Ask whether the tool lets you return to the recording, correct individual assignments and preserve corrections in the output you intend to share. Check how it handles unknown speakers and whether names require enrolment or confirmation. Review its storage and processing choices, including any separate handling of voice profiles.

If a supplier publishes an accuracy figure, inspect the task, dataset and measurement method before comparing it with another number. Word accuracy does not answer whether speaker labels are correct. A single percentage also does not tell you how the tool handled the particular interruption that changed the meaning of your meeting.

Common questions

Does diarisation tell me a speaker's real name?

Not on its own. It can distinguish anonymous voices within a recording. Connecting a label to a known person requires additional information or a separate identification process. Confirm the association before using the name as evidence.

Is Speaker 1 the same person in every recording?

Do not assume so. Anonymous labels usually describe the output of a particular recording or processing run. A product needs an additional identity mechanism to connect people across recordings; the numeral alone is not that mechanism.

Can it separate two people talking at once?

A system may indicate overlapping speaker activity, but that is not the same as producing clean, separate audio tracks or recovering every word. Check the capabilities of the specific tool and listen to the passage before trusting an attribution.

Can a transcript have correct words and incorrect speakers?

Yes. Transcription and speaker attribution solve different problems. Review both before copying a quote, assigning an action or presenting a statement as someone's position.

The useful outcome is a conversation you can navigate and verify: words, times and speaker labels that help you find the source, with uncertainty left visible where the recording does not support a confident answer.