Understanding the functional neuroanatomy of the visual system using topographic deep neural networks and spatiotemporal receptive fields

Understanding the functional neuroanatomy of the visual system using topographic deep neural networks and spatiotemporal receptive fields

🎙 Kalanit Grill-Spector 👥 4K 📅 August 12, 2026 ⏱ 56 min 👁 66 📄 original study 🧭 2026-08-15
Available in: English (current) Français

Keywords

visual systemtopographic deep neural networksspatiotemporal receptive fieldsventral streamself-supervised learning

Summary

Kalanit Grill-Spector presents her lab’s recent work on understanding the functional neuroanatomy of the visual system using topographic deep neural networks (TDNNs) and spatiotemporal receptive fields. She begins by highlighting the hierarchical and parallel organization of visual areas, with three streams (dorsal, lateral, ventral) and increasing receptive field sizes along the hierarchy. She then introduces a method to measure spatiotemporal receptive fields using fMRI and encoding models, revealing that both spatial and temporal integration windows increase systematically from V1 to higher areas, along with compressive nonlinearities. The core of the talk focuses on TDNNs, which are trained with a self-supervised task (SimCLR) and a spatial loss that encourages nearby units to have correlated responses, mimicking cortical wiring constraints. The TDNN successfully reproduces V1-like orientation, spatial frequency, and color maps with pinwheel structures, as well as category-selective clusters in VTC, matching human data. Removing the spatial loss eliminates topography, while self-organizing maps fail to match functional representations. Finally, she discusses whether the same principles can explain the organization into streams, testing the classic hypothesis of task-specific optimization versus a unified self-supervised learning with spatial constraints. The talk concludes by suggesting that a single set of principles—self-supervised learning and wiring minimization—can explain multiple scales of visual cortex organization.

206 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides significant value by introducing a novel computational model (TDNN) that unifies functional and topographic organization of the visual cortex across scales. The argumentation is strong, systematically comparing TDNN outputs to empirical data from multiple labs and using quantitative metrics like smoothness and representational similarity. The inclusion of control conditions (task-only, untrained, self-organizing maps) strengthens the causal claims about the role of spatial constraints. The presentation is well-structured, with clear hypotheses and rigorous validation.

Scientific Rigor, Source Quality, Title Accuracy

The talk is scientifically rigorous, with methods and results clearly described and grounded in established literature. The speaker cites specific studies (e.g., Yamins et al., 2014; Nauhaus, Roe) and uses publicly available datasets (ImageNet, NSD). The title accurately reflects the content, focusing on understanding visual system organization through topographic deep neural networks and spatiotemporal receptive fields. The presentation includes quantitative comparisons and validation against empirical data, enhancing its credibility.

160 words

Title / Content Match

The title accurately reflects the content, which focuses on understanding visual system organization through topographic deep neural networks and spatiotemporal receptive fields.

Quality & Reliability

9/10

The talk presents original research from a leading expert, with methods and results clearly described. The work is published in peer-reviewed venues and builds on established literature. The presentation is rigorous, with quantitative comparisons and validation against empirical data.

Key Moments

Cited Sources

  • Semir Zeki, 'The Visual Image in the Mind and the Brain' — Referenced as an influential article from Scientific American (1992) that inspired the speaker.
  • Yamins et al. (2014) — Cited for showing that DNNs trained on categorization predict ventral stream responses.
  • Nauhaus et al. — Cited for empirical data on orientation and spatial frequency maps in macaque V1.
  • Roe et al. — Cited for empirical data on color maps in V1.
  • SimCLR — Mentioned as the self-supervised learning algorithm used in the TDNN.
  • NSD dataset — Mentioned as a dataset with human fMRI responses to stimuli used for comparison.

Concurring Sources

  • Yamins et al. (2014) — Supports the predictive power of DNNs for ventral stream responses.
  • Nauhaus et al. — Empirical data on V1 maps consistent with TDNN outputs.
  • Roe et al. — Empirical data on color maps consistent with TDNN outputs.

Dissenting Sources

  • Self-organizing maps (SOMs) — SOMs produce overly smooth maps and fail to match functional representations, contrasting with TDNNs.

Contribution & Novelties

The talk presents a novel computational framework (TDNN) that unifies functional and topographic organization of the visual cortex, addressing a gap in standard DNNs. It demonstrates that self-supervised learning with a spatial constraint can reproduce multiple scales of organization, from V1 maps to VTC category clusters. This provides a unified principle for understanding cortical maps.

Pour aller plus loin :

104 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a well-rounded presentation with strong information content, technical depth, and reliability. The talk excels in providing novel insights and rigorous validation.

Reliability 9/10

💬 No comments were provided for analysis.