Agentic Financial Reasoning with Knowledge Graphs and LLMs

Agentic Financial Reasoning with Knowledge Graphs and LLMs

🎙 Abhinav Arun 👥 5K 📅 August 11, 2026 ⏱ 31 min 👁 21 📄 expert opinion 🧭 2026-08-15
Available in: English (current) Français

Keywords

knowledge graphsLLMmulti-hop QAretrievalfinance

Summary

The talk addresses the challenges of building reliable AI agents for financial reasoning. The speaker argues that finance requires more than just powerful LLMs; it needs structured evidence and knowledge graphs to ground reasoning. He presents FinReflectKG, a large-scale open-source financial knowledge graph built from SEC 10-K filings, and a benchmark called FinReflectKG-MultiHop for multi-hop question answering. The benchmark includes questions spanning intra-document, inter-year, and cross-company reasoning. Experiments with multiple reasoning LLMs and six evidence protocols show that KG-linked evidence improves correctness by 20% and reduces token usage by over 70% compared to text-window and semantic retrieval baselines. The talk concludes with a discussion on building agentic systems that orchestrate retrieval, reasoning, and verification, emphasizing an evaluation-driven development approach.

119 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into the importance of structured evidence for LLM reasoning in finance. The argument is well-supported by empirical results from the speaker’s research, including a large-scale benchmark and experiments with multiple models. The speaker effectively demonstrates that evidence organization is a critical factor, often more important than model scaling. The discussion of token efficiency and cost trade-offs is particularly useful. The argumentation is solid, though it would benefit from more discussion of potential limitations and alternative approaches.

Scientific Rigor, Source Quality, Title Accuracy

The talk is scientifically rigorous, referencing the speaker’s own published work and the FinReflectKG dataset. The methodology is clearly described, and the results are presented with quantitative evidence. The title accurately reflects the content. The speaker does not cite external sources beyond his own work, but the research appears to be well-conducted. The talk is based on a single perspective, but the evidence is compelling.

161 words

Title / Content Match

The title accurately reflects the content, focusing on agentic reasoning in finance using knowledge graphs and LLMs.

Quality & Reliability

8/10

The talk is grounded in the speaker's own research, with references to published papers and a large-scale benchmark (FinReflectKG-MultiHop). The methodology is clearly explained, and the results are presented with quantitative evidence. However, the talk is a single perspective and lacks external validation or discussion of limitations in depth.

Key Moments

Cited Sources

  • FinReflectKG paper and dataset — Mentioned as the paper published around August last year on building and evaluating knowledge graphs at scale.
  • FinReflectKG-MultiHop benchmark — Introduced as the large-scale benchmark for multi-hop QA in finance.

Concurring Sources

Dissenting Sources

Contribution & Novelties

The talk presents original research on the use of knowledge graphs for financial reasoning, introducing the FinReflectKG-MultiHop benchmark and demonstrating that structured evidence significantly improves LLM performance and efficiency. The findings challenge the assumption that reasoning LLMs alone can compensate for poor evidence organization.

Pour aller plus loin :

  • Knowledge Graph — Foundational concept for representing interconnected information.
  • Retrieval-Augmented Generation (RAG) — The framework that motivates the use of external knowledge in LLMs.
  • Multi-hop Question Answering — The task addressed in the benchmark.

83 words

Radar Profile

The radar profile shows high scores in information quantity, quality, and reliability, with a slightly lower technical level. This indicates a well-substantiated talk with strong empirical evidence, though the technical depth may be moderate for a specialized audience.

Reliability 8/10

💬 No comments were provided for analysis.