Coding with AI and Vibe Coding

Coding with AI and Vibe Coding

🎙 Minh Trinh 👥 356 📅 November 20, 2025 ⏱ 45 min 👁 44 📄 expert opinion 🧭 2026-08-15
Available in: English (current) Français

Keywords

vibe codingLLMcoding benchmarksAI toolsbest practices

Summary

This presentation by Minh Trinh explores the landscape of AI-assisted coding, focusing on large language models (LLMs) and the emerging practice of ‘vibe coding’. The speaker begins by explaining how LLMs are trained on code and their capabilities in generation, debugging, and refactoring. He then introduces key evaluation metrics like Pass@1 and Pass@K, and discusses major benchmarks including MBPP, SWE-Bench, HumanEval, LiveCodeBench, and CODEOSO. The talk covers notable coding models such as OpenAI Codex, Code Llama, DeepSeek Coder, and Claude Code, along with leaderboards and issues like benchmark contamination. The second half is dedicated to vibe coding, a term coined by Andrej Karpathy, where developers rely heavily on AI to generate code from natural language prompts without deeply reviewing the output. The speaker outlines best practices, risks, and limitations, emphasizing the need for rigorous review in production settings. He also touches on security concerns like ‘slop squatting’ and technical debt. The presentation concludes with a demo of building an app using Gemini, and discusses future directions for coding AI.

169 words

Critical Evaluation

Value of the Information & Strength of the Argument

The presentation offers valuable insights into the current state of AI coding tools and the practical aspects of vibe coding. It provides a comprehensive overview of benchmarks and models, which is useful for understanding the landscape. The argumentation is based on the author’s experience and observations, but lacks rigorous evidence or citations. The discussion of best practices and risks is practical and well-structured, though some claims could be better substantiated.

Scientific Rigor, Source Quality, Title Accuracy

The scientific rigor is moderate. The speaker references well-known benchmarks and models, but does not provide formal citations or links to specific papers. The quality of sources is acceptable for a general audience, but the lack of verifiable references weakens the overall reliability. The title accurately reflects the content, covering both AI coding and vibe coding. No comments were provided for analysis.

147 words

Title / Content Match

The title accurately reflects the content, covering both AI-assisted coding and the specific practice of vibe coding.

Quality & Reliability

7/10

The presentation is based on the author's expertise and includes references to well-known benchmarks and models, but lacks formal citations and verification of claims. It provides a broad overview with practical insights, but some information may be outdated or oversimplified.

Key Moments

Cited Sources

  • rodeo.ai — Mentioned in the video description as a resource for AI books and courses.

Concurring Sources

  • SWE-bench — The benchmark is discussed in the video and its official website provides detailed information.
  • HumanEval — The benchmark is mentioned in the video; the GitHub repository contains the dataset.
  • LiveCodeBench — The benchmark is discussed in the video; the official website provides details.

Contribution & Novelties

The presentation provides a practical overview of AI coding tools and the emerging practice of vibe coding, highlighting both opportunities and risks. It synthesizes information from various benchmarks and models, offering a useful starting point for developers. The discussion of best practices and security concerns adds practical value.

Pour aller plus loin :

  • SWE-bench — Official website for the SWE-bench benchmark, providing details and leaderboard.
  • HumanEval — GitHub repository for the HumanEval benchmark, including code and data.
  • LiveCodeBench — Official website for LiveCodeBench, a contamination-resistant benchmark.
  • Andrej Karpathy’s tweet on vibe coding — Original tweet introducing the term ‘vibe coding’.

100 words

Radar Profile

The radar profile shows high scores in quantity of information and technical level, indicating a comprehensive and technical presentation. However, the lower score in reliability suggests that the content could benefit from more rigorous sourcing and verification.

Reliability 6/10