Sam Roweis: Automatic Speech Processing By Inference in Generative Models

Sam Roweis: Automatic Speech Processing By Inference in Generative Models

🎙 Sam Roweis 👥 4K 📅 December 4, 2025 ⏱ 68 min 👁 44 📄 lecture 🧭 2026-08-16
Available in: English (current) Français

Keywords

speech processinggenerative modelsinferencesource separationvector quantization

Summary

In this lecture, Sam Roweis presents a general approach to low-level speech processing tasks using inference in generative models. He emphasizes that inference can be viewed as an automatic signal processing algorithm derived from probability and statistics. The key idea is to construct a simple generative model for the data, incorporate domain knowledge, and use inference to perform tasks like denoising and source separation. He introduces the max-VQ model, where each source is represented by a vector quantizer, and the observed spectrogram is modeled as the elementwise maximum of the sources’ contributions. This model leverages the sparsity and redundancy of speech in the time-frequency domain, where the spectrogram of a mixture is approximately the elementwise maximum of the individual spectrograms. He demonstrates that even a very simple generative model can yield surprisingly good results for denoising and single-microphone separation. He also discusses the trade-offs between working in the spectral domain versus the time domain, highlighting the importance of looking at the raw signal to understand its structure. The talk includes practical demonstrations and insights into the philosophy of using generative models for signal processing.

184 words

Critical Evaluation

Value of the Information & Strength of the Argument

The value of the information lies in its demonstration that simple generative models combined with inference can effectively solve complex speech processing tasks. The argumentation is solid, building from observations about speech sparsity and redundancy to a concrete model and results. The speaker provides clear reasoning and addresses potential criticisms, such as the unrealistic nature of the generative model, by showing that it still yields useful inferences. The demonstrations and examples strengthen the credibility of the approach.

Scientific Rigor, Source Quality, Title Accuracy

The talk is scientifically rigorous, with clear explanations of the model and its assumptions. The speaker cites relevant prior work, such as Roger Moore’s observation about the max approximation, and mentions the NoiseX database for experiments. The title accurately reflects the content. The presentation is well-structured and the technical details are appropriately explained.

145 words

Title / Content Match

The title accurately reflects the content, which focuses on automatic speech processing through inference in generative models.

Quality & Reliability

8/10

The talk presents a well-founded approach to speech processing using generative models and inference, with clear explanations and demonstrations. The methodology is based on established principles in signal processing and machine learning, and the speaker is a recognized expert. However, the content is from 2004 and may not reflect the latest state of the art.

Key Moments

Cited Sources

  • NoiseX database — Mentioned as the source of noise data for denoising experiments.

Concurring Sources

  • Roger Moore's technical report on the max approximation — The talk cites Roger Moore's observation that the spectrogram of a mixture is approximately the elementwise maximum of the sources.

Contribution & Novelties

The talk presents a novel perspective on speech processing by framing inference in generative models as an automatic signal processing algorithm. The max-VQ model is a simple yet effective approach for source separation and denoising, demonstrating that even weak generative models can be useful. The emphasis on the max approximation in the spectral domain is a key insight. The talk also encourages exploration of time-domain processing, which is often overlooked.

Pour aller plus loin :

125 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a well-rounded and reliable presentation. The talk is technically deep, provides substantial information, and is scientifically rigorous, making it a valuable resource for those interested in speech processing and generative models.

Reliability 8/10