JournalAI Data

Field guide / 7

Open Yap 1K offers 1,000 hours for free. The public sample is 8.9 hours

Open Yap 1K fills a real full-duplex speech gap, but its downloadable sample and request-only corpus have different licenses, evidence, and withdrawal rules.

Sep 6, 20267By ISH Team
Open Yap 1K offers 1,000 hours for free. The public sample is 8.9 hours
Advertisement

Open Yap 1K offers 1,000 hours for free. The public sample is 8.9 hours

Mix two speakers into one audio track and a useful part of the conversation disappears. The words remain, but it becomes harder to measure who started talking, who stopped, and whether a short sound was an interruption or a backchannel. A voice assistant trained only on orderly turns may know what to say and still feel clumsy when a person says "mm-hmm" or cuts in.

Open Yap 1K, released September 3 by The Agentic Data Company, was collected for this full-duplex problem. The company reports 1,000 hours of English conversation in 1,602 sessions from 239 speakers. Each speaker has an independent track recorded at 48 kHz and 16-bit PCM. Friends, relatives, partners, and colleagues talk without assigned topics, preserving overlap, laughter, backchannels, and interruptions.

Most readers will first encounter a much smaller release. The Hugging Face repository contains a hand-picked sample of 8.9 hours, 16 conversations, and eight speakers under CC-BY-4.0. The full 1,000-hour corpus is also available at no charge, but recipients must request it and accept a separate Data Use Agreement. It is free for commercial use. It is not a redistributable public download.

Separate channels preserve the timing

The ICASSP 2026 HumDial study treats interruption, overlap, feedback, and turn negotiation as explicit evaluation targets. Once two voices have been summed, these events are harder to isolate.

Open Yap keeps both tracks on one timeline. Across the full corpus, the publisher reports a median turn-taking gap of 580 milliseconds and median overlap of 8.3% of voiced time. At the 95th percentile, conversations reach 20 turns per minute and 20.9% overlap. These crowded exchanges test whether a system can listen while speaking.

The collection method encourages informal talk. One speaker invites somebody they know through the company's calling app. The reported relationship mix is 70.9% friends, 10.8% colleagues, 9.8% romantic partners, and 8.5% family. This is a chosen population, not a random sample of English speakers. The trade is broader representativeness for the untidy timing found between familiar people.

The sample is for inspection, not generalization

The downloadable sample is enough to test the loader and inspect the format. It has separate tracks, word-level transcripts, speaker metadata, conversation metrics, and audio-quality fields. It cannot validate claims about all 1,000 hours.

The dataset card says the sample was hand-picked and its distribution does not generalize to the corpus. Eight of the 32 tracks have no energy above 8 kHz because the speakers used Bluetooth headsets, although the containers are 48 kHz. Deepgram Nova-3 generated the transcripts, which were not human-verified. Names and overlapping speech are expected error cases.

A paper using those 16 conversations should call it a sample experiment. Until a team receives and measures the larger delivery, the 1,000-hour composition and aggregate quality figures remain publisher-reported.

The problem resembles the one in our Hindi ASR leaderboard analysis. A dataset name cannot explain a score. Readers need the unit, speaker mix, access path, and measurement details.

The full corpus has its own rules

Applicants identify themselves, describe what they are building, and explain how they will use the audio. The company reviews the request. Approved recipients receive a portal account and either a direct download or delivery to their S3 bucket.

The Data Use Agreement v1 grants a worldwide, non-exclusive, non-transferable, royalty-free license for commercial or research work. Recipients may train, fine-tune, evaluate, deploy, license, and commercialize models and outputs.

They may not redistribute the corpus, place it on another model hub, attempt to identify a speaker, or create a voice replica identifiable as somebody in the data. The agreement treats the recordings as confidential personal data. It requires access controls, encryption, a record of stored copies, and notice to the company within 72 hours of unauthorized access or loss.

Those terms shape reproducibility. Research partners cannot pass around one downloaded copy. Peers need approved access of their own. Teams may publish aggregate statistics and research results, while publishing dataset audio requires written permission from the company.

A useful experiment package can still include the dataset version, agreement version, access date, manifest hashes, split conversation IDs, preprocessing code, and permitted derived metrics. A second approved team can then reconstruct the run without receiving the voices from the first team.

Withdrawal does not recall prior deliveries

The company says speakers registered, explicitly consented before recording, and were paid. Demographics are self-reported. Stable identifiers are pseudonymous, while names, contact details, and account identifiers are excluded. A human reviewer rates language proficiency, accent classification, naturalness, and expressivity. An LLM screen checks transcripts for personal information and policy violations; flagged conversations are excluded.

The agreement nevertheless calls the records pseudonymized rather than anonymous. It also says the company does not warrant that every identifying detail was removed. Informal calls can contain family, health, work, or travel details even when account fields are stripped away.

Clause 12 gives withdrawal a limited effect. When a speaker leaves the platform, the company stops collecting from that person and stops including the recordings in future deliveries. Existing recipients do not have to delete prior copies, stop using them, or retrain a derived model. Their license over an existing delivery is perpetual. Clause 13 separately allows the company to request deletion, but a speaker's withdrawal alone does not trigger it.

People should see that rule in ordinary language before recording a family conversation. Dataset users should also record their delivery version because a later build may omit a speaker who remains in an earlier copy.

Our article on data-worker provenance argued that model lineage should include the people behind a dataset. Open Yap documents payment, explicit consent, and self-reported demographics. The retention rule belongs in that lineage as well.

Five records to keep

Before the first training or evaluation run:

  1. State whether the work used the 8.9-hour sample or an approved full-corpus delivery.
  2. Pin the dataset version, agreement version, file hashes, and conversation split.
  3. Read effective_bandwidth_hz instead of inferring usable bandwidth from the 48 kHz container.
  4. Check transcript errors around names, overlap, laughter, and short backchannels for the intended task.
  5. Store raw audio under the required controls and publish only what the agreement permits.

An inference layer such as api.ish.chat can give a team one interface while it compares models. It cannot carry the dataset license, consent record, or split definition. Those stay with the training and evaluation pipeline.

Open Yap 1K may be exactly the speech data a full-duplex project needs. Report it as one of two things: the 8.9-hour CC-BY sample or a named version of the 1,000-hour controlled corpus. That single line tells another researcher whether the result can be reproduced and tells readers which rules governed the recorded voices.

#Open Yap 1K#speech datasets#full-duplex AI#data licensing#voice privacy
Advertisement