HAI Seminar with Erik Altman: Synthetic Data Sets for the Financial Industry

HAI Seminar with Erik Altman: Synthetic Data Sets for the Financial Industry

🎙 Erik Altman 👥 34K 📅 May 21, 2025 ⏱ 72 min 👁 754 📄 expert opinion 🧭 2026-08-06
Available in: English (current) Français

Keywords

synthetic datafinancial fraudmoney launderingagent-based modelingdata labeling

Summary

Erik Altman, a research scientist at IBM, presents IBM’s Synthetic Data Sets (SDS) for the financial industry. The data sets are designed to address the scarcity and privacy issues of real financial data, providing fully synthetic data with accurate labels for fraud and criminal activity. The data is generated using an agent-based virtual world, not LLMs, creating a synthetic population with detailed attributes and referential integrity across multiple tables. The motivation includes the high rate of undetected money laundering and scams, and the limitations of real data. SDS aims to enable better AI models for detecting financial crimes, with potential benefits for banks and insurance companies. The presentation covers the data’s structure, generation process, and use cases, emphasizing the advantages of accurate labeling and global scope. Challenges include ensuring realism and avoiding biases. The talk concludes with a Q&A session.

140 words

Critical Evaluation

The seminar provides a comprehensive overview of IBM’s Synthetic Data Sets, highlighting their potential to address critical challenges in financial AI. The speaker, Erik Altman, is a credible expert with a strong background in computer architecture and AI, which lends authority to the presentation. The content is well-organized, starting with the ‘what’ and ‘why’ before delving into the ‘how’ and use cases. The emphasis on the limitations of real financial data—privacy, scarcity, and poor labeling—is compelling and well-articulated. The use of agent-based modeling instead of LLMs is a notable technical choice, and the speaker explains its advantages, such as better control over data generation and labeling. However, the presentation is largely descriptive and lacks technical depth; for instance, the specifics of the agent-based model, the validation of data realism, and the performance of models trained on SDS are not detailed. The speaker mentions that real data has high false-positive rates in fraud detection, but does not provide quantitative comparisons. The sources cited are minimal, with no specific references to academic papers or external reports, which limits the scientific rigor. The adéquation between title and content is strong, as the seminar directly addresses synthetic data sets for finance. The Q&A session, though not transcribed, likely provided additional insights. Overall, the seminar is informative and relevant, but it would benefit from more technical details and empirical evidence to strengthen its scientific contribution.

230 words

Title / Content Match

The title accurately reflects the content, which focuses on synthetic data sets for the financial industry.

Quality & Reliability

8/10

The seminar presents a well-structured overview of IBM's Synthetic Data Sets, with clear explanations of methodology and motivations. The speaker is a recognized expert, and the content is grounded in practical applications. However, the presentation is largely descriptive and lacks detailed technical validation or peer-reviewed references, limiting its scientific depth.

Key Moments

Cited Sources

  • IBM Research — IBM's research organization, where the speaker works and the Synthetic Data Sets were developed.

Concurring Sources

Dissenting Sources

  • No specific discordant sources found — The seminar did not mention any conflicting sources or viewpoints.

Contribution & Novelties

The seminar presents IBM’s Synthetic Data Sets as a novel approach to generating realistic financial data with accurate labels, addressing the critical issue of data scarcity and privacy in the financial industry. The use of agent-based modeling rather than LLMs is a distinctive methodological choice, offering potential advantages in control and interpretability. The presentation highlights the potential to improve AI models for fraud detection and other financial applications.

Pour aller plus loin :

  • Agent-based model — Relevant to the methodology used for generating synthetic data.
  • Synthetic data — Provides background on synthetic data generation and its applications.
  • Money laundering — Context for the financial crime detection use case.

108 words

Radar Profile

The radar profile shows high scores in information quantity and quality, with moderate technical depth and reliability. This indicates a well-informed presentation that is accessible but could benefit from more technical specifics and empirical evidence.

Reliability 8/10