CLSP Summer Program: Closing Presentations

CLSP Summer Program: Closing Presentations

🎙 Center for Language & Speech Processing (CLSP), JHU 👥 4K 📅 August 8, 2026 ⏱ 217 min 👁 222 📄 original study 🧭 2026-08-16
Available in: English (current) Français

Keywords

multimodal LLMencoder evaluationspecialist encodertokenizationCLSP

Summary

This video captures the closing presentations of the CLSP Summer Program at Johns Hopkins University. The main presentation, led by David Harwath, focuses on the work of the ‘omnimodal encoders’ team, which investigates multimodal large language models (LLMs). The team addresses three key problems: encoder evaluation, specialist encoders, and non-text tokenization. For encoder evaluation, they propose lightweight evaluation methods to predict the performance of encoders when integrated into LLMs, avoiding expensive training. They present initial results showing poor correlation between lightweight and heavyweight evaluations, and discuss the pitfalls of p-hacking. For specialist encoders, they explore joint training of multiple modalities and domains to create more versatile encoders. For tokenization, they investigate learning tokenization schemes for continuous inputs like audio and video to optimize performance and reduce token usage. The presentation includes detailed methodology, results, and a cautionary tale about running ablations too late. The session concludes with a discussion period.

150 words

Critical Evaluation

Value of the Information & Strength of the Argument

The presentation provides valuable insights into the challenges of building multimodal LLMs, particularly the ad hoc nature of current approaches. The team’s systematic identification of three core problems is well-argued, and they propose concrete research directions. The discussion of lightweight vs. heavyweight evaluation is particularly valuable, highlighting the need for efficient evaluation methods. The team is transparent about the limitations of their results, including the poor correlation and the issue of p-hacking, which strengthens the credibility of their argumentation. However, the presentation is primarily a progress report, and the results are preliminary, so the argumentation is more about identifying problems and proposing approaches than providing definitive solutions.

116 words

Title / Content Match

The title accurately reflects the content: a recording of the closing presentations of the CLSP Summer Program.

Quality & Reliability

7/10

Presentation of original research by a team of experts, with transparent discussion of methodology and limitations. However, results are preliminary and not peer-reviewed.

Key Moments

Cited Sources

  • MSAB benchmark — Mentioned as a source of lightweight evaluation tasks for audio.
  • MMA Pro benchmark — Used for heavyweight evaluation of multimodal LLMs.
  • Whisper — Mentioned as an example of a pre-trained audio encoder.
  • HuBERT — Mentioned as an example of a pre-trained audio encoder.
  • CLIP — Mentioned as an example of a pre-trained vision encoder.
  • DINO — Mentioned as an example of a pre-trained vision encoder.

Concurring Sources

Dissenting Sources

  • No specific discordant sources mentioned — The presentation does not discuss conflicting sources.

Contribution & Novelties

The presentation offers a systematic analysis of three open problems in multimodal LLM development: encoder evaluation, specialist encoders, and non-text tokenization. It proposes lightweight evaluation methods to predict encoder performance, explores joint training for more versatile encoders, and investigates learned tokenization schemes. The team’s honest reporting of negative results and the cautionary tale about p-hacking provide valuable methodological insights.

Pour aller plus loin :

143 words

Radar Profile

The radar profile shows high scores in technical level and information quantity, reflecting the advanced and detailed nature of the presentation. Quality of information is also high, but reliability is slightly lower due to preliminary results. The overall profile indicates a technically strong but not yet fully validated research presentation.

Reliability 7/10