Leveraging Large Speech Language Models as Evaluators for Expressive Speech

Leveraging Large Speech Language Models as Evaluators for Expressive Speech

🎙 Bismarck Odoom 👥 4K 📅 March 27, 2026 ⏱ 27 min 👁 93 📄 original study 🧭 2026-08-16
Available in: English (current) Français

Keywords

SLMexpressive speechevaluationfine-tuningemotion

Summary

The presentation introduces a method to use large speech language models (SLMs) as automatic evaluators for expressive speech attributes such as emotion, gender, emotional intensity, valence, dominance, arousal, accent, and speak rate. The speaker, Bismarck Odoom, explains the architecture of SLMs, which combine a speech encoder, a compressor, and a projector to integrate speech embeddings into a text LLM. He compares zero-shot and fine-tuned performance of models like Qwen and Qwen2 against SSL baselines on datasets like MSP-Podcast and RAVDESS. Results show that fine-tuning SLMs on just a thousand examples yields competitive or superior performance, especially for continuous attributes like arousal and valence. The speaker also discusses challenges such as data contamination and the need for cross-dataset validation. The talk concludes with future work on unified speech-text evaluation and quality assessment.

131 words

Critical Evaluation

Value of the Information & Strength of the Argument

The presentation provides valuable insights into using SLMs for expressive speech evaluation, a novel application. The argumentation is solid, supported by experimental results comparing multiple models and baselines. The speaker clearly explains the methodology and acknowledges limitations, such as potential data contamination and the need for cross-dataset studies. The discussion with the audience adds depth, addressing questions about label sets and model training. Overall, the value lies in demonstrating the potential of SLMs as scalable alternatives to human evaluation, with promising results even with limited fine-tuning data.

Scientific Rigor, Source Quality, Title Accuracy

The scientific rigor is adequate: the speaker presents original research with clear experimental setup and comparisons. However, the lack of detailed statistical analysis (e.g., confidence intervals) and the absence of cross-dataset validation weaken the conclusions. The sources cited are primarily the datasets used (MSP-Podcast, RAVDESS) and the SLM models (Qwen, Qwen2), but no external references are provided in the description. The title accurately reflects the content, and the presentation is well-structured. The discussion highlights potential issues, such as data contamination, which the speaker acknowledges.

186 words

Title / Content Match

The title accurately reflects the content, focusing on using SLMs for evaluating expressive speech.

Quality & Reliability

7/10

The presentation is based on original research with clear methodology, but lacks detailed statistical analysis and external validation. The speaker acknowledges limitations and open questions, indicating scientific honesty.

Key Moments

Cited Sources

Concurring Sources

  • MSP-Podcast dataset — Used for emotion and continuous attribute labels.
  • RAVDESS dataset — Used for emotion recognition tasks.

Contribution & Novelties

The work proposes a novel application of SLMs as automatic evaluators for expressive speech, demonstrating that fine-tuning on a small amount of data can yield competitive performance. This could reduce reliance on expensive human evaluations. The comparison with SSL baselines and the analysis of continuous attributes provide new insights.

Pour aller plus loin :

82 words

Radar Profile

The radar profile shows balanced scores across all dimensions, with slightly higher quality of information and technical level, indicating a solid but not exceptional presentation. The relatively lower quantity of information suggests the talk could have provided more details or examples.

Reliability 7/10

💬 No comments were provided for analysis.