Ep.# 170: How ChatGPT Is Used at Work, GDPval Benchmark, AI Workslop, ChatGPT Pulse, & Meta Vibes

Ep.# 170: How ChatGPT Is Used at Work, GDPval Benchmark, AI Workslop, ChatGPT Pulse, & Meta Vibes

🎙 Paul Roetzer and Mike Kaput 👥 31K 📅 September 30, 2025 ⏱ 78 min 👁 2K 📄 news review 🧭 2026-08-16
Available in: English (current) Français

Keywords

ChatGPTGDPvalAI adoptionAI benchmarksAI news

Summary

In this episode of The Artificial Intelligence Show, hosts Paul Roetzer and Mike Kaput discuss the latest developments in AI, focusing on OpenAI’s new research and products. They begin by analyzing a report on ChatGPT usage at work, which reveals that 28% of US workers use ChatGPT for their jobs, with adoption highest among younger and more educated employees. The report highlights that most users stick to basic features, while advanced capabilities like deep research are used by a minority of power users. Next, they examine OpenAI’s GDPval benchmark, a new evaluation framework that tests AI on real-world knowledge work tasks across 44 occupations. The benchmark shows that frontier models like GPT-5 and Claude Opus can produce work rated equal to or better than human experts nearly half the time, at 100 times lower cost and speed. The hosts then discuss the concept of ‘AI workslop’, a term for low-quality AI-generated content, and its implications for content creation. They also cover ChatGPT Pulse, a new feature for tracking AI trends, and Meta Vibes, a new AI model from Meta. The episode includes updates on AI and jobs, with a focus on the need for AI literacy in the workforce, and an interview with the founder of Mercor, an AI recruitment platform. Throughout, the hosts emphasize the importance of staying informed and adapting to the rapid changes in AI.

228 words

Critical Evaluation

Value of the Information & Strength of the Argument

The episode provides valuable insights into the current state of AI adoption and its economic implications. The hosts effectively argue that AI is already transforming work, citing concrete data from OpenAI’s reports and the GDPval benchmark. They balance optimism with caution, acknowledging potential risks and the need for responsible implementation. The discussion is well-structured, with each topic building on the previous one, and the hosts provide practical advice for professionals to leverage AI in their careers.

Scientific Rigor, Source Quality, Title Accuracy

The hosts demonstrate scientific rigor by referencing specific reports and studies, including OpenAI’s ChatGPT usage report and the GDPval benchmark. They also mention external sources like the Federal Reserve Bank of St. Louis for GDP data. The title accurately reflects the content, and the episode is well-organized with clear timestamps. The hosts critically evaluate the information, noting potential biases in OpenAI’s research and the limitations of the GDPval benchmark. Overall, the sources are credible and the analysis is balanced.

170 words

Title / Content Match

The title accurately reflects the main topics covered in the episode, including ChatGPT usage at work, the GDPval benchmark, AI workslop, ChatGPT Pulse, and Meta Vibes.

Quality & Reliability

8/10

The hosts provide a balanced review of recent AI news, citing specific reports and studies from OpenAI and other sources. They include critical analysis and acknowledge limitations, such as the potential bias in OpenAI's own research. The discussion is grounded in data and references, though some claims are anecdotal.

Chapters

Cited Sources

Concurring Sources

  • OpenAI GDPval Benchmark — The benchmark discussed in the episode.
  • ChatGPT Usage and Adoption Patterns at Work — The report on ChatGPT usage at work.

Dissenting Sources

  • Criticism of OpenAI's self-reported data — Some critics argue that OpenAI's reports may be biased due to their vested interest in promoting AI adoption. The hosts acknowledge this potential bias but do not provide specific counter-sources.

Contribution & Novelties

The episode provides a comprehensive overview of recent AI developments, particularly OpenAI’s efforts to measure and promote AI adoption in the workplace. The discussion of the GDPval benchmark is particularly novel, as it represents a shift from traditional AI benchmarks to evaluating real-world economic impact. The hosts also introduce the concept of ‘AI workslop’ and its implications for content quality. The episode offers practical advice for professionals to leverage AI, emphasizing the importance of AI literacy and the need to teach others.

Pour aller plus loin :

  • GDPval benchmark — Official OpenAI page for the GDPval benchmark.
  • ChatGPT usage at work report — OpenAI’s report on ChatGPT usage patterns.
  • AI workslop definition — Explanation of the term ‘workslop’.
  • Meta Vibes — Meta’s official blog post about the Vibes model.
  • ChatGPT Pulse — OpenAI’s page for ChatGPT Pulse.

137 words

Radar Profile

The radar chart shows a balanced profile with high scores in information quantity, quality, and reliability, but a slightly lower score in technical depth, indicating that the content is accessible to a general audience while still being informative.

Reliability 8/10

💬 No comments were provided for analysis.