Sergei Gukov: AI and Mathematics

Sergei Gukov: AI and Mathematics

🎙 Sergei Gukov 👥 474 📅 November 25, 2025 ⏱ 38 min 👁 114 📄 expert opinion 🧭 2026-08-03
Available in: English (current) Français

Keywords

AImathematical reasoningformal proofLeanhard problems

Summary

Sergei Gukov, a mathematician, discusses the intersection of AI and mathematics, focusing on the potential of AI to solve hard mathematical problems. He begins with the historical anecdote of Clever Hans, a horse that appeared to do arithmetic but was actually responding to subtle cues, drawing a parallel to modern AI models that may rely on biases in training data. He then demonstrates that even advanced models like GPT-4 can make simple arithmetic errors, highlighting the gap between current AI and human mathematical reasoning. Gukov outlines a progression of mathematical difficulty, from elementary problems to millennium prize problems, and notes that AI has made progress in the lower levels but remains far from research-level mathematics. He argues that solving truly hard problems will require new qualitative features, not just scaling. He suggests that many mathematical problems can be formulated as games, and that the hardness of a problem is often measured by how long it has remained open. He discusses the formalization of mathematics in languages like Lean, which turns proofs into executable code, and suggests that proving theorems is essentially pathfinding. He concludes by inviting the audience to consider whether the notion of hardness is the same for computers and humans, and what would happen if AI solved all open problems.

212 words

Critical Evaluation

The talk provides a thoughtful and accessible overview of the challenges and opportunities in using AI for mathematical reasoning. Gukov, a respected mathematician, offers a balanced perspective, acknowledging both the progress and the significant limitations of current AI systems. His use of the Clever Hans anecdote is effective in illustrating the potential for AI to rely on spurious correlations rather than genuine reasoning. The demonstration of GPT-4’s failure on a simple arithmetic problem serves as a concrete and compelling example of the current gap. The talk’s strength lies in its clear articulation of the progression of mathematical difficulty and the identification of key factors, such as the length of logical paths, that distinguish easy from hard problems. Gukov’s suggestion that many mathematical problems can be framed as games is insightful and aligns with recent developments in AI, such as AlphaGo. However, the talk is largely opinion and speculation, with no formal citations or data to support the claims. The discussion of ‘hardness’ is somewhat subjective, relying on the time a problem has remained open as a proxy. The talk would benefit from more concrete examples of AI systems tackling mathematical problems and a more rigorous analysis of the potential pathways to achieving research-level AI. The adéquation between the title and content is good, as the talk directly addresses the role of AI in mathematics. Overall, the talk is informative and thought-provoking, but it is more of an expert opinion than a rigorous scientific analysis.

244 words

Title / Content Match

The title accurately reflects the content, which discusses the potential and challenges of AI in mathematics.

Quality & Reliability

7/10

The speaker is a recognized mathematician, and the talk presents a coherent argument about AI and mathematical reasoning. However, it is largely opinion and speculation, with no formal citations or data.

Key Moments

Contribution & Novelties

The talk offers a perspective on the potential of AI in mathematics, emphasizing the need for new qualitative features beyond scaling. It frames mathematical proof as pathfinding in formal systems like Lean, and suggests that hardness is related to the length of logical paths. The talk also raises the question of whether AI could solve millennium problems and what that would mean for the field.

Pour aller plus loin :

111 words

Radar Profile

The radar chart shows a balanced profile with moderate scores across all dimensions, indicating a talk that is informative but not highly technical or data-driven. The highest score is in 'qualite_information' and 'fiabilite_globale', reflecting the speaker's expertise, while 'quantite_information' is slightly lower due to the speculative nature.

Reliability 7/10