AddSubAddSub
Guide

Speaker Diarization: Who Said What in Your Recordings

August 2026 · 5 min read

The short answer: Speaker diarization separates audio by who is speaking and labels each line (S1, S2, S3…) — and AddSub does it live as you watch. Interviews, podcasts, and meetings become transcripts where you always know who said what.

Why speaker labels matter

A plain transcript of a two-person interview is a wall of text — you can’t tell who’s asking and who’s answering. Diarization solves that by assigning each segment a persistent speaker identity. When you play back a recorded interview, a podcast with several hosts, or a team meeting, AddSub colors each line by speaker (S1, S2, S3…) so the conversation reads naturally.

How AddSub’s diarization works

  1. Voice separation — a segmentation model finds where speakers change.
  2. Voice embeddings — each segment’s voice is fingerprinted.
  3. Clustering — similar voices are grouped into speaker identities.
  4. Persistent labeling — identities are matched across the session, so S1 stays S1.

The key detail: AddSub’s labels are commit-once — once a speaker is identified, their label doesn’t flip mid-session, so transcripts stay stable over long recordings.

Getting the best accuracy

  • Keep background noise low — clean two-speaker audio gives near-perfect separation.
  • Give it ~30 seconds — labels stabilize as voices are clustered; the first few lines may converge.
  • Distinct voices help — very similar voices (e.g. two same-sex speakers on phone audio) are harder to separate.
  • Works live and offline — the real-time stream and exported transcripts both carry speaker badges.

Label every speaker — free for 7 days

No payment needed to start. Download for macOS.

Experience real-time subtitles with AddSub

Free 7-day trial. 100% on-device AI transcription and live translation for macOS.