Session 5: Data in the Age of Generative AI

Session 5: Data in the Age of Generative AI

🎙 Stanford HAI 👥 34K 📅 October 30, 2025 ⏱ 41 min 👁 305 📄 expert opinion 🧭 2026-08-06
Available in: English (current) Français

Keywords

synthetic datatraining datadata attributioncopyrightLLM

Summary

The talk, presented by James Zou and colleagues at Stanford HAI, explores the critical role of data in generative AI. It begins by highlighting the rapid consumption of human-generated text data by LLMs, projecting that models will soon exhaust available data and increasingly rely on synthetic data. The presentation is structured around two main questions: data created by AI (synthetic data) and data used by AI (training data). The first part includes a survey quantifying the prevalence of LLM-generated text in scientific papers, consumer complaints, press releases, job postings, and UN documents, using a statistical model based on word frequency differences. Dan Ho then discusses using synthetic data to improve access to sensitive health data, proposing a tiered access system with synthetic data for prototyping. Tatsu Hashimoto addresses the basic science of synthetic data, questioning its mathematical foundations and potential for model collapse. The second half focuses on data attribution, with projects on understanding how training data influences model creativity and behavior, including a study on diffusion models and a method for attributing text generation to specific training data. The talk concludes with a Q&A session.

186 words

Critical Evaluation

The talk provides a comprehensive and insightful overview of the challenges and opportunities surrounding data in generative AI. The speakers are credible experts from Stanford, and the content is grounded in ongoing research projects supported by the Hoffman-Yee grant. The quantitative estimates of LLM-generated text in various domains are intriguing and highlight the pervasive influence of AI, though the methodology is only briefly described, limiting the ability to assess its robustness. The discussion of synthetic data for public impact is particularly valuable, addressing the real-world bottleneck of data access in sensitive domains like healthcare. The proposal to combine synthetic data with tiered access is pragmatic and could significantly expand research capabilities. The section on the basic science of synthetic data raises important questions about its mathematical underpinnings and potential pitfalls, such as model collapse, but does not provide definitive answers, reflecting the nascent stage of this research. The data attribution projects are innovative and have practical implications for copyright and model transparency. However, the talk is more of a research overview than a deep dive, and some claims lack detailed evidence. The Q&A section is not transcribed, so the audience’s reactions are unknown. Overall, the talk is informative and thought-provoking, but it would benefit from more detailed methodological explanations and references to published work.

214 words

Title / Content Match

The title accurately reflects the content, which focuses on the role of data in generative AI, covering both synthetic data and training data.

Quality & Reliability

8/10

The talk is presented by Stanford faculty and researchers, providing a credible overview of ongoing research on data in generative AI. It includes quantitative estimates and references to specific projects, but lacks detailed methodological transparency and peer-reviewed sources in the description.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

The talk provides a novel framework for understanding the dual role of data in generative AI: as both a product and a resource. It presents original research on quantifying the prevalence of LLM-generated text in various domains, which is a significant contribution to understanding the impact of AI on information ecosystems. The discussion on using synthetic data to democratize access to sensitive datasets is innovative and has practical implications for research and policy. The data attribution work offers new methods for tracing model outputs to training data, which is crucial for addressing copyright and accountability issues.

Pour aller plus loin :

185 words

Radar Profile

The radar profile shows high scores in information quantity, quality, and reliability, reflecting the talk's comprehensive coverage and credible sources. The technical level is moderately high, indicating that the content is accessible to a general audience but still provides depth for experts.

Reliability 8/10