Basis Conversations 1500: 1,500 hours of multilingual, multi-party conversation
2,645 speakers meeting, sharing, arguing, laughing, trolling across 22 languages. Each conversation includes up to four simultaneous speakers, each with channel-separated, 48 kHz audio.
Blog + Samples · Get the dataset
What's in the release
| Conversation audio | 1,502 hours |
| Conversations | 1,907, delivered as 2,396 segments |
| Unique speakers | 2,645 across 33 countries |
| Languages | 22 |
| Audio | 48 kHz, 16-bit mono FLAC; one track per speaker |
| Human annotations | About 100 hours; 113,096 judgments from 7,362 labelers |
Hours are measured once on the conversation timeline, not added up across speaker tracks.
Multi-party. A conversation seats between two and four people at a time; participants come and go over the course of the conversation. The median conversation lasts 33 minutes.
Multilingual. English, Spanish, Japanese, Korean, Georgian, Mingrelian, German, French, Italian, Hebrew, Russian, Hindi, Arabic, Ukrainian, Dutch, Xhosa, Zulu, Portuguese, Chinese, Turkish, Polish, and Kazakh. The dataset card breaks down hours, conversations, speakers, and annotated hours by language.
Track-separated. Each participant's device records its own microphone. Every track of a segment starts at the same instant and has the same length, so a timestamp indexes the same moment across all speakers. Audio is provided as recorded, with no denoising or loudness normalization.
The subtleties of conversation
About 100 hours are densely labeled with human judgments on subtleties that models struggle with:
- Was this backchannel sympathetic or frustrated?
- Was this silence awkward or turn-holding?
- Was this utterance directed at one speaker or everyone?
- Was this overlap cooperative or competitive?
- Was the speaker correcting themselves, restarting, or abandoning a thought?
The release includes 23,751 human-judged moments in 148 conversations across 14 languages, with about five votes per moment. Labels cover backchannels, laughter, overlaps and interruptions, pauses, repairs, self-repairs, and who an utterance was addressed to.
We use waveform analysis, offline audio models, and LLMs judging transcripts to identify candidate events, then ask people to categorize them. Human judgments and machine labels are provided separately. For laughter and backchannels, we only include machine labels that people confirmed; for the other events, use the human-judged moments if you want a checked set.
You can listen to labeled examples before downloading.
How we collected it
All conversations were recorded on our custom consumer applications, where participants are paid for their contributions. Participants were matched with strangers who speak the same language and talked about whatever they wanted for as long as they liked.
Recording environments and microphones vary across speakers. We filter out conversations with very poor audio defects while retaining some noisy, real-life conditions for robust training and evaluation.
Contributors consent to the recording and licensing of their contributions for AI research and commercial development. Spoken personal details are muted where identified, and cuts and muted spans are documented. Speaker identifiers are pseudonymous and stable across the release; demographics are self-reported.
Start with the metadata
Transcripts, annotations, and speaker metadata can be inspected without downloading audio:
from huggingface_hub import snapshot_download
local_dir = snapshot_download(
repo_id="basis-ai/basis-conversations-1500",
repo_type="dataset",
allow_patterns=[
"*.jsonl", "*.json", "*.md",
"manifests/*", "LICENSE", "CHECKSUMS.sha256",
],
)
To add Spanish audio, include "es/audio/*/*" in allow_patterns. The full repository is approximately 253 GB. The README also includes examples for 🤗 Datasets and Lhotse.
Transcripts are automatically generated and can be inaccurate, especially for less-supported languages. They are provided for finding and filtering moments of interest; we do not recommend using the supplied transcripts as training targets without further checking. Languages are unevenly represented.
Why we're releasing it
Progress in real-time human-AI interaction depends heavily on data that captures people listening, responding, and speaking to one another across languages. We're making 1,500 hours of conversational data available so that researchers and builders worldwide can study human interactions and develop better models: understanding nuances of human conversations, turn-taking, emotional understanding, humor, and more.
Basis has over one million contributors submitting audio, video, annotations, ratings, and transcriptions. We'll keep releasing interaction data to cover the weird, ambiguous, natural modes of human interaction.
Download Basis Conversations 1500. Free for commercial and research use under the dataset's license. For custom collections, reach us at data@withbasis.co.
