[Video Response] What Cloudflare's code mode misses about MCP and tool calling

[Video Response] What Cloudflare's code mode misses about MCP and tool calling

🎙 Yannic Kilcher 👥 329K 📅 October 19, 2025 ⏱ 13 min 👁 10K 📄 expert opinion 🧭 2026-08-16
Available in: English (current) Français

Keywords

MCPtool callingcode modeLLMspeculative decoding

Summary

Yannic Kilcher responds to Theo’s video and Cloudflare’s article on ‘code mode’, which proposes converting MCP tools into TypeScript APIs for LLMs to call via code generation. He agrees that this approach leverages LLMs’ pre-training on code, making it more natural than traditional tool calling. However, he critiques the article’s claim that code mode shines for multi-step tool calls by avoiding intermediate LLM context. He argues that in real-world scenarios, intermediate outputs are often non-deterministic and require reasoning, so pre-composing all calls into code may fail. He suggests a hybrid approach: execute code with speculative tool calls, then have the LLM validate all intermediate outputs at the end. He also reminds viewers that MCP is just a standard for exposing APIs, not magic. The video is a concise, thoughtful commentary on the limitations and potential of code mode.

138 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video provides valuable insights into the practical limitations of code mode for complex tool orchestration. Kilcher’s argument that intermediate outputs often require reasoning is well-articulated and supported by a relatable example. He also introduces a novel idea of ‘speculative tool calling’ with validation, which adds original value. The argumentation is logical and balanced, acknowledging the benefits of code mode while highlighting its potential pitfalls.

Scientific Rigor, Source Quality, Title Accuracy

The video is rigorous in its references, citing the Cloudflare article and Theo’s video directly. The title accurately reflects the content. The argumentation is based on reasoning and practical experience rather than empirical data, but it is coherent and well-structured. The video does not overstate claims and appropriately caveats its points.

131 words

Title / Content Match

The title accurately reflects the content, which is a response to Cloudflare's code mode article and Theo's video, focusing on the missed aspects of MCP and tool calling.

Quality & Reliability

8/10

The video is a well-reasoned expert commentary by a recognized AI researcher. It clearly references the Cloudflare article and Theo's video, and provides a balanced critique. The arguments are logical and grounded in practical experience, though they are not backed by empirical data or formal experiments.

Key Moments

Cited Sources

Concurring Sources

Dissenting Sources

  • Theo's video: MCP is the wrong abstraction — Theo's video argues that MCP is the wrong abstraction, while Kilcher suggests that MCP is fine but code mode may miss some nuances. This is a point of disagreement.

External References

Contribution & Novelties

The video offers a critical perspective on Cloudflare’s code mode, highlighting a potential flaw in its approach to multi-step tool calls. It introduces the concept of ‘speculative tool calling’ with validation, which is an original idea not present in the original article. This adds value by proposing a hybrid method that could combine the benefits of code mode with the flexibility of traditional tool calling.

Pour aller plus loin :

  • Model Context Protocol (MCP) — Official documentation for MCP, the standard discussed in the video.
  • Speculative Decoding — Research paper on speculative decoding, which inspired the proposed ‘speculative tool calling’.
  • Toolformer — A paper on teaching language models to use tools, relevant to the discussion of tool calling.

118 words

Radar Profile

The radar profile shows high scores in quality and reliability, with moderate scores in quantity and technical depth. This indicates a well-reasoned expert opinion with solid argumentation, but not an exhaustive or highly technical analysis.

Reliability 8/10