Best Model for RAG? GPT-4o vs Claude 3.5 vs Gemini Flash 2.0 (n8n Experiment Results)

Best Model for RAG? GPT-4o vs Claude 3.5 vs Gemini Flash 2.0 (n8n Experiment Results)

🎙 Nate Herk 👥 964K 📅 January 30, 2025 ⏱ 18 min 👁 13K 📄 expert opinion 🧭 2026-08-28
Available in: English (current) Français

Keywords

RAGLLMn8nGPT-4oClaude 3.5

Summary

The video presents a casual experiment comparing three large language models—OpenAI’s GPT-4o, Anthropic’s Claude 3.5 Sonnet, and Google’s Gemini Flash 2.0—for use in Retrieval-Augmented Generation (RAG) agents built with n8n. The creator explains the basics of RAG, then outlines a seven-test methodology: information recall, query understanding, response coherence and completeness, speed, context window management, handling conflicting information, and source attribution. Each test uses the same prompt across all models, and responses are graded by GPT-4o (with access to the source PDF) to maintain consistency. Results show Claude 3.5 leading with an average score of 8.6, GPT-4o at 7.7, and Gemini Flash 2.0 at 6.9. Notable findings include Gemini’s superior speed (6.7 seconds vs. 11 and 21 seconds) and Claude’s strong performance in coherence and context management. The creator acknowledges the experiment’s limitations, such as using an OpenAI model to grade responses and the potential for bias. The video concludes with practical advice: the best model depends on the use case, and testing multiple models is recommended.

166 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video provides a practical, hands-on comparison of LLMs for RAG, which is valuable for practitioners building AI agents. The creator demonstrates a clear methodology, albeit informal, and transparently shares the results. The argumentation is straightforward: the experiment is designed to be consistent, and the creator acknowledges its imperfections. The value lies in the practical insights, such as Gemini’s speed advantage and Claude’s strength in generating coherent, well-structured responses. However, the lack of rigorous statistical analysis and the reliance on a single grader (GPT-4o) weaken the argument’s robustness. The creator’s reasoning for grading is explained, but the subjective nature of the evaluations is a limitation.

Scientific Rigor, Source Quality, Title Accuracy

The video does not cite external scientific sources; it relies on the creator’s own experiment. The description includes links to n8n (with an affiliate link), Skool communities, and social media, but these are not sources for the content. The title accurately reflects the content, as it is indeed an experiment comparing models for RAG. The methodology is described in detail, but the lack of peer-reviewed references or external validation reduces the scientific rigor. The creator’s transparency about the experiment’s limitations is a positive aspect, but the overall reliability is limited by the informal nature of the test.

217 words

Title / Content Match

The title accurately reflects the content: a comparative experiment of three LLMs for RAG, with results presented.

Quality & Reliability

5/10

The video presents a casual, non-scientific experiment with acknowledged methodological limitations, such as using a single grader (GPT-4o) and inconsistent grading criteria. The creator is transparent about these limitations, but the lack of rigorous controls and the absence of peer-reviewed sources reduce the overall reliability.

Chapters

Cited Sources

Concurring Sources

  • n8n - Workflow Automation — The tool used to build the RAG agents in the experiment.

Contribution & Novelties

The video offers a practical, hands-on comparison of three major LLMs for RAG, which is valuable for practitioners. It provides a replicable methodology (though informal) and highlights performance differences in speed, coherence, and context handling. The main novelty is the direct comparison in a real-world n8n workflow, offering insights that are often missing from theoretical discussions.

Pour aller plus loin :

123 words

Radar Profile

The radar profile shows moderate scores across all dimensions, with a slight dip in reliability. This reflects the video's practical but informal nature: it provides useful information and a decent technical level, but the lack of rigorous methodology and external validation lowers its overall reliability.

Reliability 4/10