BioML Seminar 4.2 - Kenny Workman on Agents for Real-World Spatial Biology Data

BioML Seminar 4.2 - Kenny Workman on Agents for Real-World Spatial Biology Data

🎙 Kenny Workman 👥 14K 📅 May 2, 2026 ⏱ 72 min 👁 357 📄 original study 🧭 2026-08-16
Available in: English (current) Français

Keywords

SpatialBenchspatial transcriptomicsAI agentsbenchmarkbioinformatics

Summary

Kenny Workman, CTO of LatchBio, presents SpatialBench, a benchmark for evaluating AI agents on real-world spatial biology data analysis. He begins by contextualizing the exponential growth of biological data, driven by technologies like single-cell sequencing and spatial transcriptomics. He explains the complexity of spatial data, which combines imaging and molecular counts, and the immature analysis ecosystem. The core of the talk is the design of SpatialBench: 146 verifiable problems derived from real workflows, covering five spatial technologies and seven task categories. He emphasizes the importance of verifiability, durability, and anti-shortcut properties in benchmark design, illustrating with examples of good and bad test questions. He details the process of decomposing workflows into verifiable chunks, using domain experts and quality control. He shows how grading functions are crafted, using examples like clustering and marker gene identification. He reports that frontier models achieve only 20-38% accuracy, highlighting the gap between current AI capabilities and the demands of real biological analysis. The talk concludes with implications for the future of AI in biology and the potential for agents to automate complex analysis tasks.

179 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into the practical challenges of applying AI agents to scientific data analysis. The argumentation is strong, grounded in the speaker’s direct experience building a benchmark and a company around biological data infrastructure. The speaker clearly explains the rationale behind benchmark design choices, such as verifiability and durability, and supports claims with concrete examples. The presentation of benchmark results is honest, acknowledging the low accuracy of current models. The talk also offers a broader perspective on the future of AI in biology, arguing that agents will increasingly handle complex analysis tasks. The argumentation is persuasive and well-supported, though it is primarily based on the speaker’s own work and may lack independent validation.

Scientific Rigor, Source Quality, Title Accuracy

The talk demonstrates scientific rigor in the design and evaluation of SpatialBench. The methodology is transparent, with clear criteria for problem selection and grading. The speaker acknowledges limitations, such as the difficulty of ensuring durability and the potential for shortcuts. The quality of sources is high, as the talk references real technologies and workflows, and the speaker is a practitioner in the field. The title accurately reflects the content, focusing on agents for spatial biology data. The talk does not rely on external citations but rather presents original work, which is appropriate for a seminar. Overall, the scientific rigor is commendable, though the lack of external references limits the ability to verify claims independently.

245 words

Title / Content Match

The title accurately reflects the content: a seminar on agents for spatial biology data, presented by Kenny Workman.

Quality & Reliability

8/10

The talk presents a novel benchmark (SpatialBench) with a clear methodology, verifiable problems, and quality control. The speaker is CTO of LatchBio, providing practical context. Limitations are acknowledged, and the approach is transparent.

Key Moments

Cited Sources

  • SpatialBench: A Benchmark for Agents on Real-World Spatial Biology Data — The benchmark introduced in this talk, not yet published.

Concurring Sources

Contribution & Novelties

The talk introduces SpatialBench, a novel benchmark for evaluating AI agents on real-world spatial biology data analysis. It addresses a critical gap in existing benchmarks by focusing on verifiable, workflow-derived problems that require genuine data interaction. The benchmark’s design principles—verifiability, durability, and anti-shortcut—provide a framework for creating robust evaluations in complex scientific domains. The talk also offers practical insights into the challenges of building such benchmarks, including the need to decompose workflows and craft grading functions. This contributes to the growing field of AI for scientific discovery, highlighting the gap between current model capabilities and the demands of real-world analysis.

Pour aller plus loin :

135 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a well-rounded and reliable presentation. The talk excels in providing substantial information, maintaining high technical depth, and demonstrating strong reliability through transparent methodology.

Reliability 8/10