Building Modern Databases with the FDAP Stack

Building Modern Databases with the FDAP Stack

🎙 Andrew Lamb & Olimpiu Pop 👥 1.1M 📅 November 24, 2025 ⏱ 29 min 👁 2K 📄 expert opinion 🧭 2026-08-02
Available in: English (current) Français

Keywords

FDAPcolumnar storageApache ArrowApache ParquetDataFusion

Summary

In this GOTO Unscripted interview, Andrew Lamb, staff engineer at InfluxData and PMC member of Apache DataFusion and Apache Arrow, discusses the FDAP stack (Flight, DataFusion, Arrow, Parquet) as a modular approach to building modern databases. He explains the shift from row-based to columnar storage, driven by the need to handle increasing data volumes and leverage modern hardware. Apache Arrow provides a standardized in-memory columnar format, enabling efficient data exchange and processing without serialization overhead. Apache Parquet serves as a columnar file format for persistent storage, offering compression and efficient reads. DataFusion is a query engine that processes data in Arrow format, and Flight is a network protocol for efficient data transfer. The discussion highlights how these components can be assembled to build a database like InfluxDB 3.0, saving years of development time. The interview also touches on the future with Apache Iceberg for interoperability and the trend towards modular data systems.

152 words

Critical Evaluation

The interview provides a valuable, expert-level overview of modern database architecture, specifically the FDAP stack. Andrew Lamb’s credentials as a staff engineer at InfluxData and PMC member of Apache DataFusion and Arrow lend significant authority to the discussion. The content is technically accurate and reflects current best practices in the field, such as columnar storage and the use of standardized components. The argumentation is coherent, moving from the historical context of database evolution to the specific components of the FDAP stack and their roles. The discussion is well-structured, with clear explanations of each technology and its purpose. The sources cited, including the InfluxData blog post and the CIDR paper, are relevant and credible, though the interview itself is not a formal scientific presentation. The title accurately reflects the content, and the interview delivers on its promise. The main strength is the practical insight from someone who has built a database using these components, providing a real-world perspective. The main limitation is the interview format, which may lack the depth of a technical paper, but it serves as an excellent introduction to the topic. Overall, the information is reliable and well-presented, making it a valuable resource for developers and data engineers interested in modern database design.

205 words

Title / Content Match

The title accurately reflects the content, which focuses on building modern databases using the FDAP stack.

Quality & Reliability

8/10

High-quality expert discussion by a staff engineer and PMC member of Apache DataFusion and Arrow, providing practical insights into modern database architecture. The claims are grounded in real-world experience and reference open-source projects, though the format is an interview rather than a peer-reviewed presentation.

Chapters

Cited Sources

Concurring Sources

External References

Contribution & Novelties

The interview provides a clear, practitioner-oriented explanation of the FDAP stack, emphasizing the benefits of assembling databases from standardized open-source components. It highlights how this approach reduces development time and improves interoperability. The discussion of columnar storage and the roles of Arrow, Parquet, DataFusion, and Flight is insightful for developers considering building their own data systems.

Pour aller plus loin :

  • Apache Arrow — Official documentation and resources for Apache Arrow, the in-memory columnar format.
  • Apache Parquet — Official documentation for Apache Parquet, the columnar file format.
  • Apache DataFusion — Official documentation for Apache DataFusion, the query engine.
  • Apache Iceberg — Official documentation for Apache Iceberg, a table format for large analytic datasets.

113 words

Radar Profile

The radar profile shows high scores in information quantity, quality, and reliability, with a slightly lower technical level, indicating a well-balanced presentation that is accessible yet informative.

Reliability 8/10