Tutorial: Incentives for Collaborative Learning and Data Sharing, Part II

Tutorial: Incentives for Collaborative Learning and Data Sharing, Part II

🎙 Sai Praneeth Karimireddy 👥 75K 📅 February 17, 2026 ⏱ 87 min 👁 362 📄 tutorial 🧭 2026-08-06
Available in: English (current) Français

Keywords

data valuationShapley valueleave-one-outincentivescollaborative learning

Summary

This tutorial, part of the Federated and Collaborative Learning Boot Camp at the Simons Institute, focuses on data valuation and incentives for data sharing. The speaker, Sai Praneeth Karimireddy, introduces the concept of data-centric AI and the market failure for individual data, motivating the need for data intermediaries. He discusses the fundamental question of attributing value to individual data contributors in a machine learning pipeline. The tutorial covers the simplest method, leave-one-out error, and highlights the challenge of stochasticity in modern deep learning, where the value function is a random variable. The speaker illustrates the issue with examples and engages in a detailed Q&A with the audience, discussing the interpretation of variance, the dependence between data points, and the practical implications. The talk aims to bridge theory and practice, emphasizing the need for scalable, efficient, and justifiable attribution methods. The speaker also mentions a real-world scenario involving YouTube and music copyright, illustrating the complexity of attribution in generative models. The tutorial is interactive, with audience questions shaping the discussion, and concludes with open questions for future research.

177 words

Critical Evaluation

The tutorial provides a solid introduction to data valuation, a key topic in collaborative and federated learning. The speaker, Sai Praneeth Karimireddy, is an expert in the field, and his presentation is clear and well-structured. He starts with the motivation for data-centric AI and the market failure for individual data, which sets the stage for the need for data valuation. The core of the talk focuses on the leave-one-out method and its limitations in the context of stochastic deep learning. The speaker effectively illustrates the issue of variance in influence scores with a histogram and discusses the implications for defining contribution. The interactive Q&A session is particularly valuable, as it addresses nuanced questions about the interpretation of variance, the dependence between data points, and the practical challenges of applying these methods. The speaker acknowledges the limitations of current methods and highlights open questions, which is intellectually honest. However, the tutorial is introductory and does not delve into advanced techniques or recent research in depth. The hypothetical YouTube scenario, while illustrative, is not based on verified facts, which the speaker acknowledges. The sources cited are limited to the Simons Institute talk page, which provides context but not detailed references. Overall, the tutorial is informative and thought-provoking, but it would benefit from more concrete examples and references to recent literature. The title accurately reflects the content, and the presentation is suitable for a technical audience familiar with machine learning concepts.

238 words

Title / Content Match

The title accurately reflects the content, which focuses on incentives for collaborative learning and data sharing, specifically data valuation methods.

Quality & Reliability

8/10

The tutorial is given by an expert researcher in the field, with a formal mathematical approach to data valuation. The content is well-structured, discusses fundamental concepts and open questions, and includes interactive Q&A. However, it is a tutorial, not a peer-reviewed study, and some claims are based on hypothetical scenarios.

Key Moments

Cited Sources

Concurring Sources

  • Simons Institute talk page — The talk page provides official information about the tutorial, aligning with the content presented.

Contribution & Novelties

The tutorial provides a comprehensive overview of data valuation methods, highlighting the challenges posed by stochasticity in deep learning. It emphasizes the need for scalable, efficient, and justifiable attribution methods, and discusses open questions in the field. The interactive Q&A adds practical insights.

Pour aller plus loin :

102 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a well-rounded tutorial with strong information content, quality, technical depth, and reliability. The weakest point is the limited number of cited sources, but the interactive format and expert presentation compensate.

Reliability 8/10