Multispeaker Datasets for Speaker Diarization

Multispeaker diarization datasets with timestamped who-said-what speaker labels. Built for training and evaluating production speaker diarization systems.

  • Who-Said-What: Timestamped Speaker Labels
  • Conversational Data from 2-5 Speakers
  • Train Models to Separate and Identify Speakers
  • Ideal for Call Center & Meeting Transcription

Dataset information

Creator
YPAI
Formats
audio/wav, audio/flac

Check availability and fit

Review the current catalog, then request the license, provenance, metadata, and delivery details for the dataset you are evaluating.

Get Diarization Data