CLSP Summer Program: Plenary Speaker and Weekly Progress Report

CLSP Summer Program: Plenary Speaker and Weekly Progress Report

🎙 Shvit Banga 👥 4K 📅 July 25, 2026 ⏱ 80 min 👁 152 📄 expert opinion 🧭 2026-08-16
Available in: English (current) Français

Keywords

ASRTTSbenchmarkhuman evaluationdata collection

Summary

The talk, given by Shvit Banga at the CLSP Summer Program, argues that large-scale human-in-the-loop research is feasible and not as costly as commonly perceived. Banga, founder of a non-profit and a company called Jos, shares his journey from a small Himalayan town to building AI benchmarks. He presents three projects: the Voice of India benchmark for ASR, a TTS leaderboard, and a third project (not detailed). The Voice of India benchmark involved 36,000 speakers, 15 languages, and 536 hours of audio, collected in 90 days at a cost of $5,000. It covers nine axes including geography, age, gender, vocabulary, devices, acoustic environment, speech type, speech speed, and multiple valid transcripts. The talk emphasizes the importance of including multiple valid transcripts to accurately measure model performance. The TTS leaderboard aims to cover more languages and domains than existing benchmarks. Banga concludes that the infrastructure for large-scale human studies exists and encourages researchers to embrace such challenges.

156 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into the practical challenges and solutions for building large-scale human-in-the-loop benchmarks. The speaker’s argument that such studies are feasible and cost-effective is supported by concrete examples and data from his projects. However, the argumentation is largely based on anecdotal evidence and lacks rigorous statistical analysis or comparison with existing benchmarks. The speaker’s enthusiasm and practical experience add credibility, but the lack of peer-reviewed validation weakens the overall argument.

82 words

Title / Content Match

The title is generic and does not reflect the specific content about large-scale human-in-the-loop benchmarks.

Quality & Reliability

7/10

The speaker provides detailed insights into large-scale human-in-the-loop data collection and evaluation for ASR and TTS, based on practical experience. However, the talk is largely anecdotal and lacks peer-reviewed validation or detailed methodological transparency.

Key Moments

Cited Sources

  • Voice of India Benchmark — Mentioned as an ASR benchmark presented at Interspeech 2026.
  • Jos (company) — Mentioned as the speaker's company, which monetizes data sets.

Concurring Sources

Contribution & Novelties

The talk provides a compelling case for the feasibility of large-scale human-in-the-loop research in AI, with specific examples from ASR and TTS benchmarks. The emphasis on multiple valid transcripts and demographic representation is a novel contribution. The speaker’s practical experience and cost estimates offer valuable guidance for researchers.

Pour aller plus loin :

  • Interspeech — Conference where the Voice of India benchmark was presented.
  • Common Voice — A crowdsourced dataset for speech recognition, similar in spirit.
  • TTS Leaderboard — Existing TTS leaderboard mentioned in the talk.

86 words

Radar Profile

The radar profile shows high scores in information quantity and quality, with moderate technical depth and reliability. This indicates a talk rich in practical insights but with limited formal rigor.

Reliability 6/10

💬 No comments were provided for analysis.