Exploring High-Bandwidth Flash (HBF) for Modern LLM-Serving Systems

Exploring High-Bandwidth Flash (HBF) for Modern LLM-Serving Systems

🎙 Prof. Jisung Park 👥 64K 📅 July 25, 2026 ⏱ 89 min 👁 1K 📄 original study 🧭 2026-08-15
Available in: English (current) Français

Keywords

HBFLLM inferencememory capacitybandwidthGPU

Summary

The talk, presented by Prof. Jisung Park, explores the potential of High-Bandwidth Flash (HBF) as a scalable complement to DRAM-based HBM for large language model (LLM) serving systems. HBF, proposed by SanDisk in 2025, aims to provide ~20x larger capacity with comparable read bandwidth to HBM by exploiting massive parallelism in NAND flash. The speaker systematically evaluates HBF-based systems using simulation, comparing them to HBM-only baselines across various configurations and workloads. Key findings indicate that HBF can significantly reduce GPU count requirements and improve throughput, but achieving HBM-comparable read bandwidth and endurance is critical. The talk also highlights technical challenges such as write performance and endurance, and suggests that a hybrid HBM+HBF approach may be more practical. The research is based on recent large models (Llama 3 and Llama 4) and realistic workloads, providing insights for future memory system design.

140 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into an emerging memory technology and its application to LLM serving. The argumentation is solid, based on systematic simulation and comparison with state-of-the-art HBM systems. The speaker clearly explains the methodology, assumptions, and results, and addresses potential limitations such as the optimistic performance projections. The work is original and contributes to understanding the feasibility and benefits of HBF, which is crucial for guiding future hardware and system design.

Scientific Rigor, Source Quality, Title Accuracy

The talk demonstrates high scientific rigor, with a clear methodology and reliance on established simulation tools. The sources cited include the speaker’s own publications and industry projections from SanDisk, which are relevant and credible. The title accurately reflects the content, and the presentation is well-structured. The talk is based on a paper accepted for publication, indicating peer review. The speaker also acknowledges assumptions and future work, enhancing credibility.

156 words

Title / Content Match

The title accurately reflects the content, focusing on exploring HBF for LLM-serving systems.

Quality & Reliability

8/10

The talk presents original research with a systematic methodology, including simulation-based analysis and comparisons with industry projections. The speaker is an established researcher in memory systems, and the work is accepted for publication at a top venue (ITP). However, the results rely on assumptions and simulations, and the technology is still emerging, so absolute certainty is limited.

Key Moments

Cited Sources

Concurring Sources

  • SanDisk's HBF proposal — Industry proposal for HBF technology (mentioned in talk)

Contribution & Novelties

The talk provides a systematic analysis of HBF for LLM serving, offering insights into its potential benefits and challenges. It highlights that HBF can significantly reduce GPU count and improve throughput, but emphasizes the need for HBM-comparable read bandwidth and endurance enhancements. The work also suggests that a hybrid HBM+HBF approach may be more practical. This contributes to guiding future research and development in memory systems for AI.

Pour aller plus loin :

101 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a well-rounded and reliable presentation. The talk excels in information quantity and quality, with strong technical depth and credibility. The overall high scores reflect the speaker's expertise and the rigorous methodology.

Reliability 8/10

💬 Sur les 0 commentaires analysés, aucune tendance n'est disponible.