Reno Kriz, Toward Event-Centric Video Retrieval and Summarization: (SCALE 2024 and SCALE 2026

Reno Kriz, Toward Event-Centric Video Retrieval and Summarization: (SCALE 2024 and SCALE 2026

🎙 Reno Kriz 👥 4K 📅 September 1, 2026 ⏱ 76 min 👁 0 📄 expert opinion 🧭 2026-09-01
Available in: English (current) Français

Keywords

video retrievalmultimodalevent-centricSCALERAG

Summary

Reno Kriz, a research scientist at Johns Hopkins University’s Human Language Technology Center of Excellence (HLTCOE), presents an overview of two summer workshops (SCALE 2024 and SCALE 2026) focused on event-centric video retrieval and summarization. The talk begins by framing the problem: the vast amount of user-generated video content, especially raw footage from events, is difficult to retrieve using traditional methods that rely on metadata and text descriptions. Kriz contrasts this with professionally produced news videos, which are easier to process. He introduces the MultiVENT dataset, a collection of multilingual videos of events, and its extension MultiVENT 2.0, which includes 100,000 distractor videos and 2,500 queries. The SCALE 2024 workshop brought together around 50 researchers to tackle this retrieval problem using multiple modalities: vision, speech, OCR, and text. Key findings from SCALE 2024 include: extracting and using text from video (via OCR and speech transcription) significantly improves retrieval performance; LLM-based summarization of noisy OCR text is beneficial; and no single modality is sufficient—fusion of all modalities yields the best results. The talk also highlights that retrieving raw event footage is substantially harder than retrieving edited videos. SCALE 2026 builds on this by focusing on the hardest setting: raw videos. The workshop is structured in two stages: first, evaluating modality-specific technologies (audio/visual event detection, speech/audio summarization, OCR/visual frame analysis); second, a multimodal retrieval-augmented generation (RAG) task where systems must retrieve relevant segments from a large collection of raw multilingual videos and generate a coherent summary. The talk concludes with an invitation to the final readout event.

255 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into the practical challenges of multimodal video retrieval, particularly for raw, user-generated content. The speaker effectively argues that current methods, such as YouTube search, rely heavily on metadata and text, which are often unavailable or unreliable for raw footage. The presentation of SCALE 2024’s findings—that text extraction and fusion of modalities are key—is well-supported by the described experiments and results. The argumentation is solid, as the speaker grounds his claims in specific datasets and models (CLIP, Whisper, PaddleOCR) and openly discusses limitations, such as the difficulty of OCR in multilingual settings and the challenges of raw video. The talk also clearly motivates the progression to SCALE 2026, which addresses the identified gaps. However, the talk is more of a high-level overview than a detailed technical exposition, and some claims, such as the superiority of text IR models over multimodal encoders, are stated without presenting specific quantitative comparisons.

Scientific Rigor, Source Quality, Title Accuracy

The talk demonstrates a strong scientific rigor, as it is based on the speaker’s direct involvement in the SCALE workshops and the development of the MultiVENT datasets. The speaker references specific papers (e.g., the MultiVENT paper at NeurIPS 2023, the MultiVENT 2.0 paper at CVPR 2025) and models (CLIP, Whisper, PaddleOCR), which adds credibility. The title accurately reflects the content, which is a summary of SCALE 2024 and a preview of SCALE 2026. The talk is well-structured, with a clear outline and logical progression from problem definition to solutions and future directions. The speaker also engages with audience questions, providing thoughtful responses that clarify technical details. However, the talk does not provide a full list of references or links to the mentioned papers, which would be useful for further verification.

297 words

Title / Content Match

The title accurately reflects the content, which covers takeaways from SCALE 2024 and directions for SCALE 2026.

Quality & Reliability

8/10

The talk is given by a research scientist leading the SCALE workshops, providing direct insight into the design, results, and ongoing challenges of a large-scale research effort. The content is grounded in specific datasets (MultiVENT, MultiVENT 2.0) and published papers (CLIP, Whisper, PaddleOCR), and the speaker openly discusses limitations and open questions. However, the talk is a high-level overview without detailed experimental methodology or quantitative results, and some claims rely on anecdotal evidence.

Key Moments

Cited Sources

  • MultiVENT: Multilingual Videos of Events with Aligned Natural Text — Mentioned as the original dataset paper, accepted at NeurIPS 2023.
  • MultiVENT 2.0 — Mentioned as the updated dataset with 100,000 videos and 2,500 queries, accepted at CVPR 2025.
  • CLIP — Mentioned as a baseline for visual embeddings.
  • Whisper — Mentioned as the speech transcription model.
  • PaddleOCR — Mentioned as the OCR system used for text extraction.

Concurring Sources

  • MultiVENT: Multilingual Videos of Events with Aligned Natural Text — The dataset paper, which the talk is based on, supports the claims about the dataset's construction and challenges.
  • CLIP — The use of CLIP embeddings for visual retrieval is a standard approach, and the talk's findings align with its known limitations.
  • Whisper — Whisper is a well-known speech recognition model, and its use in the pipeline is consistent with its capabilities.

Dissenting Sources

  • YouTube search — The talk argues that YouTube search is insufficient for raw video retrieval, which contrasts with the assumption that YouTube's search is effective for general video retrieval.

Contribution & Novelties

The talk provides a comprehensive overview of the SCALE 2024 and SCALE 2026 workshops, highlighting key findings and future directions in event-centric video retrieval. The main novelty is the focus on raw, unedited video footage, which is a challenging and under-explored area. The talk also emphasizes the importance of multimodal fusion and the use of LLMs for summarization of noisy OCR text. The progression from MultiVENT to MultiVENT 2.0 and the introduction of a multimodal RAG task in SCALE 2026 represent significant steps forward in the field.

Pour aller plus loin :

129 words

Radar Profile

The radar profile shows high scores in information quantity, quality, and reliability, reflecting the speaker's expertise and the depth of the content. The technical level is also high, indicating a detailed discussion of methods and challenges. The overall profile suggests a well-rounded and informative talk.

Reliability 8/10