Karen Livescu: Factoring Speech into Linguistic Features

Karen Livescu: Factoring Speech into Linguistic Features

🎙 Karen Livescu 👥 4K 📅 December 14, 2025 ⏱ 69 min 👁 90 📄 lecture 🧭 2026-08-16
Available in: English (current) Français

Keywords

speech recognitionarticulatory featurespronunciation modelingdynamic Bayesian networksspeech variability

Summary

Karen Livescu presents her research on factoring speech into linguistic features for automatic speech recognition. She begins by highlighting the high variability in speech, exemplified by multiple pronunciations of the word ‘probably’. She argues that traditional phone-based pronunciation models struggle with sparse data and low coverage of conversational pronunciations. She proposes a feature-based approach using articulatory features, such as lip opening, tongue position, and voicing, which can capture partial changes and asynchrony between articulators. The model is implemented using dynamic Bayesian networks, allowing for learning and inference. She demonstrates that feature-based models can explain pronunciation variations more parsimoniously than phone-based models, with examples like the insertion of a ’t’ in ‘sense’ due to desynchronization of articulators. Experimental results show improvements in speech recognition performance, and she discusses ongoing work on context-dependent feature models and applications to other tasks.

138 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into the limitations of phone-based pronunciation models and proposes a novel feature-based approach. The argumentation is solid, supported by examples and experimental results. The speaker effectively demonstrates how articulatory features can capture pronunciation variability more accurately and parsimoniously. The use of dynamic Bayesian networks is well-motivated, and the discussion of implementation details adds credibility. However, the talk is primarily a research presentation, and the experimental evidence is not exhaustive; some claims could benefit from more extensive validation.

Scientific Rigor, Source Quality, Title Accuracy

The talk is scientifically rigorous, with clear references to prior work in articulatory phonology and speech recognition. The speaker cites her own PhD thesis and summer workshop collaborations, which are relevant and credible. The title accurately reflects the content. The presentation is well-structured and the methodology is transparent. However, as a single lecture, it does not provide a comprehensive review of the literature, and some sources are mentioned but not fully cited. The adequacy between title and content is high.

177 words

Title / Content Match

The title accurately reflects the content, which focuses on decomposing speech into articulatory features for pronunciation modeling.

Quality & Reliability

8/10

The talk is a technical lecture by a recognized researcher, based on her PhD thesis and subsequent work, with clear methodology and references to published research. The content is well-structured and grounded in established speech science and machine learning. However, it is a single presentation without peer review, and some claims rely on specific datasets and models that are not fully detailed.

Key Moments

Cited Sources

  • Feature-based pronunciation modeling (PhD thesis) — The speaker's own thesis, which forms the basis of the talk.
  • Johns Hopkins Summer Workshop on Articulatory Feature-based ASR — Mentioned as a collaboration where part of the work was done.
  • Switchboard corpus — Used for pronunciation variability statistics and experiments.

Concurring Sources

Contribution & Novelties

The talk presents a novel approach to pronunciation modeling by using articulatory features, which offers parsimony and linguistic interpretability compared to phone-based models. The implementation via dynamic Bayesian networks is a significant contribution, enabling learning and integration into standard ASR frameworks. The examples illustrate how feature asynchrony can explain common pronunciation phenomena.

Pour aller plus loin :

84 words

Radar Profile

The radar profile shows high scores in technical level and information quality, indicating a deep and reliable presentation. The quantity of information is also high, but the overall score is slightly lower due to the niche nature and lack of broad accessibility.

Reliability 8/10