What is speaker diarization?

Speaker diarization splits an audio recording into segments by speaker, so a transcript shows who spoke when. It answers "who talked" rather than "what was said." In meetings it turns a wall of text into attributed lines. Use Wispr Notetaker to get your words and your speakers right in every meeting

How speaker diarization works

Diarization runs on the audio, not the words. A system splits the recording into short segments, measures the vocal characteristics of each one, and groups segments that sound like the same person. The output is a timeline: speaker 1 from 0:00 to 0:14, speaker 2 from 0:14 to 0:31, and so on.

Those groups start out anonymous. Diarization can tell that two people are talking without knowing either name. Naming them is a separate step, and it usually comes from outside the audio: a calendar invite, a participant list, or someone in the room saying "what do you think, Dan?"

A few things make the job harder. People talk over each other. One person joins from a conference room while everyone else has their own microphone. Someone steps away and comes back with a different headset. Clean one-on-one audio is easy. Those edge cases are where diarization quality is decided.

Speaker diarization vs transcription vs speaker recognition

These three get used interchangeably, and they answer different questions. Transcription answers what was said. Diarization answers who spoke when, without needing to know who anyone is. Speaker recognition answers which known person is talking, by matching a voice against a stored profile.

TermQuestion it answersWhat it needsTypical output
TranscriptionWhat was saidAudioText, with timestamps
Speaker diarizationWho spoke whenAudioAnonymous speaker segments on a timeline
Speaker labelingWhich name goes on each segmentDiarization plus outside contextNamed speakers in a transcript
Speaker recognitionWhich known person is speakingA stored voice profileAn identity match

The difference matters when you are comparing tools. A tool can transcribe well and still hand you generic speaker labels. A tool can diarize well and still leave the names to you. Wispr Notetaker does not build voiceprints or biometric profiles of you or anyone on your call, so recognition is not how it gets names onto the page.

What diarization looks like in a real meeting

Take a three-person call about a partnership. Siobahn asks how the NVIDIA partnership is going. Emeka says they are stalled until SOC 2 compliance is done. Xiaoyan says she will talk to their Vanta rep later.

Without diarization you get one block of text with three voices in it, and no way to tell which of the three said they would talk to Vanta. With diarization you get three separate speaker turns. With names attached, you can see who said what: Xiaoyan is the one talking to Vanta, and Emeka is the one who raised SOC 2.

This is where the downstream cost shows up. A follow-up email built on a misattributed line goes to the wrong person. If you pipe meetings into Claude or ChatGPT, one wrong label travels through everything you ask afterward.

How Wispr Notetaker handles who said what

Wispr Notetaker separates speakers during the call and names them after it. It reads your calendar invite and, if you connect them, Slack and Gmail, plus your personal dictionary. Uncommon names come out spelled right because they are on the invite. Internal acronyms come out right because they are in your dictionary.

During the meeting you get a live transcript, with speakers shown as You and Them. Afterward it goes back over the audio and produces a final transcript, and that is where speakers get their names. It uses the invite, the context from those apps, and clues in the conversation itself. If it cannot identify someone, you name them once and the label applies across the whole transcript.

Nothing joins the call, so there is no bot in the participant list. Tell people you are recording anyway. Wispr Notetaker records on your device, which is why it works across Zoom, Google Meet, Teams, and a Slack huddle rather than one platform at a time.

Where diarization still struggles

Diarization struggles in four setups: shared rooms, heavy crosstalk, very short turns, and people who are never addressed by name.

  1. Shared rooms. Several people on one microphone are the hardest case
  2. Heavy crosstalk. When two people speak at once, separating one turn from the next is harder
  3. Very short turns. A one-word "agreed" carries little vocal information to group on
  4. Unnamed participants. Anyone in the room who is never addressed by name may not get a label

Practical fixes are unglamorous. Put the meeting on the calendar with everyone on the invite. Have people address each other by name.

Frequently asked questions

How accurate is speaker diarization?

It depends on the audio. Separation is easiest when each person has their own microphone and audio stream. Shared room microphones and one-word turns are where labels slip. Judge separation apart from naming: a tool can split the speakers correctly and still show generic labels until names arrive from an invite or from you.

How many speakers can diarization handle?

It depends on the system and the audio. Grouping gets harder as the number of voices climbs, and a large call where several people share one room microphone gives the system the fewest cues to work with.

Why does my meeting transcript show generic speaker labels instead of names?

That means the tool separated the speakers but has not attached names yet. In Wispr Notetaker the generic You and Them labels are expected during the call, and names appear once the final transcript is built. If someone is left unnamed, you name them once.

Does diarization work for in-person meetings?

It works, but it is the harder setup. Put the meeting on the calendar with everyone on the invite, and have people address each other by name so the labels have something to go on. Anyone in the room who is never named may not get a label.

Does Wispr Notetaker keep recordings of my meetings?

Only for a short time. Meeting audio is encrypted and kept only temporarily to create your transcript, let you resume a meeting, verify quality, and troubleshoot. After that limited period, it's automatically removed. Wispr Notetaker doesn't create voiceprints or biometric profiles of you or anyone on your call.

Share this