Speaker Diarization: Who Said What in Your Recordings
August 2026 · 5 min read
The short answer: Speaker diarization separates audio by who is speaking and labels each line (S1, S2, S3…) — and AddSub does it live as you watch. Interviews, podcasts, and meetings become transcripts where you always know who said what.
Why speaker labels matter
A plain transcript of a two-person interview is a wall of text — you can’t tell who’s asking and who’s answering. Diarization solves that by assigning each segment a persistent speaker identity. When you play back a recorded interview, a podcast with several hosts, or a team meeting, AddSub colors each line by speaker (S1, S2, S3…) so the conversation reads naturally.
How AddSub’s diarization works
- Voice separation — a segmentation model finds where speakers change.
- Voice embeddings — each segment’s voice is fingerprinted.
- Clustering — similar voices are grouped into speaker identities.
- Persistent labeling — identities are matched across the session, so S1 stays S1.
The key detail: AddSub’s labels are commit-once — once a speaker is identified, their label doesn’t flip mid-session, so transcripts stay stable over long recordings.
Getting the best accuracy
- Keep background noise low — clean two-speaker audio gives near-perfect separation.
- Give it ~30 seconds — labels stabilize as voices are clustered; the first few lines may converge.
- Distinct voices help — very similar voices (e.g. two same-sex speakers on phone audio) are harder to separate.
- Works live and offline — the real-time stream and exported transcripts both carry speaker badges.
Label every speaker — free for 7 days
No payment needed to start. Download for macOS.
Experience real-time subtitles with AddSub
Free 7-day trial. 100% on-device AI transcription and live translation for macOS.