Do Multilingual Embedding Models Really Retrieve African Language Content Well? A Yoruba Case Study

Do Multilingual Embedding Models Really Retrieve African Language Content Well? A Yoruba Case Study

🎙 Adumi Joshua 👥 278 📅 August 8, 2026 ⏱ 45 min 👁 9 📄 tutorial 🧭 2026-08-16
Available in: English (current) Français

Keywords

embeddingretrievalYorubaMIRACLNDCG

Summary

This tutorial by Adumi Joshua, a graduate student at Carnegie Mellon University Africa, demonstrates how to evaluate multilingual embedding models on the Yoruba subset of the MIRACL benchmark. The session begins by explaining the difference between embedding models and generative models, and how embeddings convert text into vectors for semantic similarity. It then introduces the MIRACL dataset, which contains Wikipedia passages and 119 queries for Yoruba. The presenter defines key metrics: NDCG@10, which measures ranking quality, and Recall@100, which measures whether relevant passages are retrieved at all. Four embedding models are compared: BGE-M3, a fine-tuned version on Nigerian languages, multilingual E5, and LABSE. The tutorial includes a worked example of cosine similarity and metric calculation. After running the evaluation, results show that base BGE-M3 and multilingual E5 perform well on NDCG@10, while the fine-tuned BGE-M3 also shows strong performance. The presenter discusses the potential for fine-tuning embedding models using contrastive learning to improve performance on low-resource languages. The session concludes with insights on how to extend this evaluation to other languages and build better retrieval systems.

176 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video provides a practical, hands-on approach to evaluating embedding models, which is valuable for practitioners working on information retrieval in low-resource languages. The argumentation is clear and logical, building from basic concepts to specific metrics and results. The presenter effectively explains the importance of NDCG and Recall and demonstrates their application on a real benchmark. The comparison of multiple models offers useful insights, though the analysis could be deepened with more discussion of why certain models perform better.

88 words

Title / Content Match

The title accurately reflects the content, which focuses on evaluating multilingual embedding models for Yoruba retrieval.

Quality & Reliability

7/10

The video is a technical tutorial demonstrating evaluation of embedding models on a benchmark. It provides clear explanations of concepts and metrics, and uses a real dataset (MIRACL). However, it lacks formal citations and the results are presented without rigorous statistical analysis.

Key Moments

Cited Sources

Concurring Sources

  • MIRACL benchmark — The benchmark used in the video, consistent with its description.

Contribution & Novelties

The video provides a practical tutorial on evaluating multilingual embedding models for a low-resource language (Yoruba), which is a valuable contribution for researchers and practitioners. It demonstrates a reproducible workflow using MIRACL and open-source models. The comparison of base and fine-tuned models offers insights into the benefits of fine-tuning for specific languages.

Pour aller plus loin :

91 words

Radar Profile

The radar profile shows balanced scores across all dimensions, indicating a well-rounded tutorial with solid information quality and technical depth, though not exceptional in any single area.

Reliability 7/10