Emulating Real-World PII with a Large-Scale Synthetic Dataset to Audit LLM Memorization

Emulating Real-World PII with a Large-Scale Synthetic Dataset to Audit LLM Memorization

🎙 Sriram Selvam, Anneswa Ghosh 👥 5K 📅 August 11, 2026 ⏱ 34 min 👁 108 📄 original study 🧭 2026-08-15
Available in: English (current) Français

Keywords

PANORAMAPIImemorizationsynthetic datasetprivacy audit

Summary

The talk by Sriram Selvam and Anneswa Ghosh from Microsoft AI introduces PANORAMA, a large-scale synthetic dataset designed to emulate real-world Personally Identifiable Information (PII) in online content. The dataset contains 384,789 samples derived from 9,674 synthetic human profiles, spanning six online modalities (social media, forum comments, online reviews, article comments, marketplace ads, and wiki-style biographies). The generation pipeline uses constrained selection and reasoning LLMs, paired with Wikipedia articles to enhance realism while maintaining synthetic integrity. The speakers detail the challenges of creating realistic synthetic data and the importance of internal consistency in profiles. They then present an audit study where they continued pre-training Mistral 7B on PANORAMA with varying repetition rates, measuring memorization via ROUGE-L F1 and soft match rate. Key findings include increased memorization with repetition, with online ads showing significantly higher memorization than social media. They also measured PII leakage, which rose from near-zero baseline to 44% after training on PANORAMA. The dataset has been widely adopted by industry and academia, highlighting its utility for privacy research and mitigation development.

173 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable contributions by addressing a critical gap in privacy research: the lack of realistic datasets with PII for studying LLM memorization. The PANORAMA dataset is a significant resource, being the largest of its kind at release. The argumentation is solid, with clear motivation, methodology, and experimental results. The speakers effectively demonstrate the dataset’s utility through a concrete audit study, showing how it can be used to measure memorization and PII leakage. They also discuss the importance of internal consistency in synthetic profiles and the use of Wikipedia scaffolding to enhance realism. The findings on content-type-specific memorization are novel and provide insights for future research. The presentation is well-structured and persuasive, though it could benefit from more details on the generation pipeline and potential limitations.

Scientific Rigor, Source Quality, Title Accuracy

The talk demonstrates scientific rigor through a systematic approach to dataset creation and evaluation. The speakers mention using entity recognition filters to remove profiles with Wikipedia contamination, ensuring the synthetic nature of the data. They also compare different LLMs for generation, selecting O3 for its balance of originality and realism. The audit study uses established metrics (ROUGE-L F1, soft match rate) and bootstrap resampling for confidence intervals. The title accurately reflects the content, and the presentation is well-aligned with the abstract. The speakers cite prior work on memorization (e.g., Nicholas Carlini’s divergence attack) and mention the lack of PII in existing datasets like The Pile and RedPajama. However, they do not provide specific citations for these claims, and the talk is not a peer-reviewed publication, so the rigor is high but not at the level of a formal paper.

281 words

Title / Content Match

The title accurately reflects the content: the talk introduces a synthetic dataset to emulate real-world PII and uses it to audit LLM memorization.

Quality & Reliability

8/10

The talk presents a novel synthetic dataset (PANORAMA) and a rigorous audit methodology, with clear experimental design and results. The speakers are from Microsoft AI, and the work is open-source. However, the presentation is a conference talk, not a peer-reviewed paper, and some details (e.g., exact training hyperparameters) are omitted.

Key Moments

Cited Sources

Concurring Sources

Dissenting Sources

  • No discordant sources found — The talk does not present conflicting evidence; it builds on prior work.

Contribution & Novelties

The talk introduces PANORAMA, a novel synthetic dataset that addresses the lack of realistic PII data for studying LLM memorization. Its key contributions include: (1) a large-scale dataset with 384,789 samples from 9,674 profiles, spanning six online modalities; (2) a generation pipeline that combines synthetic profiles with Wikipedia scaffolding to enhance realism while maintaining synthetic integrity; (3) an audit study demonstrating the dataset’s utility in measuring memorization and PII leakage, with findings on content-type-specific memorization. The dataset is open-source and has been adopted by industry and academia, filling a critical gap in privacy research.

Pour aller plus loin :

168 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a well-rounded presentation with strong information content, technical depth, and reliability. The lowest score is in technical level, suggesting it is accessible to a broad audience while still providing substantial detail.

Reliability 8/10

💬 No comments were provided for analysis.