JournalAI Research

Field guide / 5

Vaani timestamps 106,892 real noise events. Only 21.85 hours are in its verified tier

Vaani captures real overlapping noise across 58 Indian languages, but its verified and unverified timestamp tiers should not be treated as equivalent.

Sep 8, 20265By ISH Team
Vaani timestamps 106,892 real noise events. Only 21.85 hours are in its verified tier
Advertisement

Vaani timestamps 106,892 real noise events. Only 21.85 hours are in its verified tier

A dog barking halfway through a sentence poses a different problem from a noise track pasted over clean speech. The bark begins near a particular word, may overlap with traffic, and stops while the speaker carries on. Synthetic mixtures can make a signal noisy without teaching a model that sequence.

The Vaani Noise Event Timestamp Dataset records the sequence. ARTPARK and the Indian Institute of Science built it on Project Vaani field recordings. Its release-quality subset has 106,892 timestamped noise events in 72,756 speech segments, covering 58 Indian languages, 30 states, and 162 districts.

Only 21.85 hours belong to the verified timestamp tier. Another 100.32 hours has timestamps without confirmed inter-annotator agreement. Publishing the split prevents readers from assuming that the full label count has one level of certainty.

Real noise has timing

Many noisy-speech corpora start with separate recordings of clean speech and noise. A researcher selects a signal-to-noise ratio and mixes the tracks. The result is controlled and repeatable, but the noise has no natural relationship to the words, device, room, or speaker.

Vaani's source audio came from ordinary mobile devices in homes, farms, streets, and other field settings. Speech and background sounds happened together. Each annotated event has a category, a specific tag, and start and end times. Events may overlap each other and the speech.

This supports a more precise test than asking whether a clip contains traffic. Researchers can check whether a model finds the horn at the correct moment, whether recognition errors cluster around it, and whether speech enhancement removes the disturbance without cutting nearby words.

Seven categories organize the sounds: animal; vehicle and traffic; baby and child; singing and music; phone, signal, and alarm; appliance and machine; and non-speech human sounds. Non-speech human events appear in 37.8% of segments and number 37,739, yet total only 4.5 hours because coughs and lip smacks are short. Animal events cover 22.4 hours. Vehicle and traffic events cover 12.5 hours. Counting events alone would hide that contrast.

Why the release has two hour totals

The paper's release-quality subset contains 122.17 hours from 38,541 speakers. It excludes segments with structural issue flags, held-out evaluation data, synthetic data, and records outside the two timestamp tiers. Of the 72,756 included segments, 72,746 have at least one timestamped event.

The Hugging Face dataset card describes a broader current release of 90,637 segments and about 154.6 hours. Its total includes 17,884 no_timestamps segments, about 32.4 hours. The card also says another 10 hours of verified, speaker-disjoint evaluation data is held out and absent from the release.

The figures refer to different slices. Reporting only 154.6 hours could suggest that every released clip has event boundaries. Reporting only 122.17 hours overlooks category-only audio that may still help with classification. A benchmark needs to name the slice it used.

Language coverage is uneven too. Hindi contributes 83.9 hours and 47,080 segments in the paper. Telugu contributes 16.7 hours, Bengali 12.9, and Marathi 5.3. A long tail includes Chakma, Garo, and Mizo. Coverage across 58 languages does not mean balanced representation across them.

What verification means here

Freelancers marked the start, end, and type of audible events. A structural sanity check covered the full output. About 100 hours that passed entered unverified_timestamps.

For the higher-trust tier, an internal team re-timestamped at least 20 hours. A second reviewer audited a random 10% sample. One disagreement sent the batch back for rework. The released verified tier contains 11,111 segments and 21.85 hours.

Researchers can therefore choose between scale and confidence. The paper does not give a representative agreement rate for the unverified labels, so it cannot tell us their error rate. Combining the tiers without distinction would discard the clearest piece of provenance in the dataset.

Use the verified tier for evaluation and detailed error analysis. The unverified data can support training or weak supervision, but keep the tiers separate and report whether adding the lower-confidence labels changes performance. The no_timestamps subset fits tasks that require categories rather than exact boundaries.

Design a useful robustness test

One aggregate word error rate is not enough. Break results down by language, noise category, event duration, and overlap with speech. Add state or district when sample counts support it. Compare errors inside annotated spans with errors elsewhere in the clip. Include a clean or low-noise baseline, so a denoiser cannot look good simply by suppressing audio.

Train and evaluation sets should keep speakers separate. The dataset card's held-out 10-hour speaker-disjoint set follows that rule, although it is not publicly released. Any local split should document speaker separation and avoid tuning on verified test material.

Vaani uses a CC BY 4.0 license. Downloading the files still requires a Hugging Face account, acceptance of conditions, and sharing contact information. This is open licensing with gated delivery rather than an anonymous download. Read the access conditions before claiming that a pipeline is freely reproducible.

The Hindi Open ASR Leaderboard changed its metric because speech evaluation can move when scoring rules change. Vaani exposes another part of the same measurement problem: a suitable language metric still needs audio that resembles deployment. The Open Yap 1K release offers a related lesson about separating a public sample from the headline corpus size.

Vaani does not report a new model result showing that these annotations improve recognition. It releases the material needed to test that claim, along with enough provenance to show that 122 hours does not carry one confidence level. The first useful comparison is verified-only evaluation against training runs that add the unverified tier.

Primary sources

#Vaani#speech recognition#Indian languages#datasets#AI evaluation
Advertisement

Keep reading

Related stories

Browse the archive