Bhuvana Ramabhadran: Recent Advances in Audio Information Retrieval

Bhuvana Ramabhadran: Recent Advances in Audio Information Retrieval

🎙 Bhuvana Ramabhadran 👥 4K 📅 December 12, 2025 ⏱ 72 min 👁 29 📄 expert opinion 🧭 2026-08-16
Available in: English (current) Français

Keywords

spoken term detectionaudio indexingsub-word unitsOOVspeech recognition

Summary

The talk, given by Bhuvana Ramabhadran at Johns Hopkins University in 2008, presents IBM’s work on spoken term detection (STD), a task of searching for keywords in audio streams. The speaker outlines the architecture of their system, which uses an ASR system to generate word confusion networks and sub-word transcripts. The system indexes both words and sub-word units (fragments) to handle out-of-vocabulary (OOV) terms. The talk covers the process of building a fragment inventory from phonetic n-grams, the trade-offs between word and sub-word indexing, and the challenges of phonetic decoding. It also discusses the NIST 2006 evaluation and the importance of precision and recall. The speaker shares insights on pronunciation handling, the use of exact vs. fuzzy matching, and the impact of word error rate on performance. The talk concludes with a discussion of practical applications in call centers and broadcast news.

142 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into the design and implementation of a spoken term detection system, based on the speaker’s extensive experience at IBM. The argumentation is solid, with a clear logical flow from the problem definition to the system architecture and experimental findings. The speaker addresses questions from the audience, clarifying technical details and acknowledging limitations. However, the talk lacks quantitative results and comparisons with other approaches, which would strengthen the argumentation. The focus is on the system’s design choices and lessons learned, rather than on rigorous experimental validation.

99 words

Title / Content Match

The title accurately reflects the content, which focuses on recent advances in audio information retrieval, specifically spoken term detection.

Quality & Reliability

8/10

The talk is given by an expert from IBM with deep experience in speech recognition and spoken term detection. It describes a specific system architecture and evaluation results, but lacks detailed quantitative results and external citations. The content is technical and appears reliable, but is limited to the speaker's own work.

Key Moments

Cited Sources

  • CLSP Seminar page — The seminar page for this talk, providing context and possibly additional materials.

Concurring Sources

  • CLSP Seminar page — The seminar page for this talk, providing context and possibly additional materials.

Contribution & Novelties

The talk provides a detailed account of IBM’s spoken term detection system, highlighting the use of sub-word units to handle OOV terms. The approach of building a fragment inventory from pruned phonetic n-grams is a notable contribution. The speaker also discusses practical considerations such as memory footprint and speed, which are often overlooked in academic research.

Pour aller plus loin :

  • Spoken term detection — Overview of the task and its challenges.
  • NIST Spoken Term Detection Evaluation — Official NIST page for the evaluation mentioned in the talk.
  • Word confusion networks — Concept used in the system for indexing.
  • Phonetic indexing — Related technique for audio search.

107 words

Radar Profile

The radar profile shows high scores in quality of information, technical level, and reliability, with a slightly lower score in quantity of information. This indicates a technically dense and reliable talk, but with limited breadth of content.

Reliability 8/10