Comp. Arch. - Lecture 30: GPU Programming (Fall 2025)

Comp. Arch. - Lecture 30: GPU Programming (Fall 2025)

🎙 Onur Mutlu 👥 64K 📅 January 12, 2026 ⏱ 152 min 👁 2K 📄 lecture 🧭 2026-08-15
Available in: English (current) Français

Keywords

GPUCUDASIMTTensor CoresMemory Bandwidth

Summary

This lecture, part of the Computer Architecture course at ETH Zürich, focuses on GPU programming. The instructor, Prof. Onur Mutlu, begins by outlining the agenda: the typical programming structure for GPUs using CUDA/OpenCL, the bulk synchronous parallel (BSP) model, memory hierarchy and management, and performance considerations. He then reviews the evolution of GPU architectures, from the Tesla architecture to the Volta V100, highlighting the increase in stream processors and the introduction of tensor cores for deep learning. The lecture explains the SIMT execution model, warp scheduling, and the organization of GPU cores (SMs) with their lanes and functional units. It discusses the programming model, emphasizing the SPMD paradigm and the ease of programming compared to traditional SIMD. The instructor addresses bottlenecks such as PCIe transfer and memory bandwidth, comparing CPU vs. GPU design philosophies. The lecture covers the steps of offloading computation to the GPU: data transfer, kernel execution, and result retrieval. It also touches on collaborative computing and mentions recommended readings, including the CUDA programming guide and academic papers on memory-centric computing and RowHammer.

175 words

Critical Evaluation

Value of the Information & Strength of the Argument

The lecture provides substantial value by offering a comprehensive overview of GPU programming, from architectural details to programming models and performance considerations. The argumentation is solid, grounded in established concepts and supported by references to academic literature and official documentation. The instructor effectively explains complex topics like SIMT execution, warp scheduling, and memory hierarchy, making them accessible to students. The discussion of bottlenecks and trade-offs is particularly valuable, as it provides a critical perspective on GPU computing. The lecture also encourages further exploration by posing questions about tensor cores and comparisons with other accelerators, fostering deeper understanding.

Scientific Rigor, Source Quality, Title Accuracy

The lecture demonstrates high scientific rigor, with content based on well-established knowledge in computer architecture. The instructor cites recommended readings, including the CUDA programming guide and several academic papers, which are listed in the video description. The sources are credible and relevant to the topic. The title accurately reflects the content, as the lecture is indeed about GPU programming within a computer architecture course. The lecture is well-structured and follows a logical progression, enhancing its reliability as an educational resource.

192 words

Title / Content Match

The title accurately reflects the content: a lecture on GPU programming within a computer architecture course.

Quality & Reliability

9/10

Lecture by a renowned professor in computer architecture, with detailed technical content, references to academic papers and official course materials. The content is well-structured and based on established knowledge in GPU programming.

Key Moments

Cited Sources

Concurring Sources

External References

Contribution & Novelties

This lecture provides a comprehensive and up-to-date overview of GPU programming, covering both fundamental concepts and recent architectural developments such as tensor cores. It offers a balanced perspective on the trade-offs and bottlenecks in GPU computing, making it valuable for students and practitioners. The lecture also connects GPU programming to broader topics like memory-centric computing and RowHammer, providing a holistic view of modern computer architecture.

Pour aller plus loin :

114 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a well-rounded and reliable lecture. The high scores in information quantity and quality reflect the comprehensive coverage of GPU programming, while the technical level and global reliability are also strong, making it an excellent educational resource.

Reliability 9/10

💬 Sur les 0 commentaires analysés, aucune tendance n'a pu être dégagée.