Language AI in the Space Sciences: Day 3 - Session 4 - March 11, 2026

Language AI in the Space Sciences: Day 3 - Session 4 - March 11, 2026

🎙 STScI Research 👥 1K 📅 March 12, 2026 ⏱ 95 min 👁 396 📄 conference presentation 🧭 2026-08-18
Available in: English (current) Français

Keywords

LLMclassificationADSJWSTHST

Summary

This session from the Language AI in the Space Sciences workshop features a talk by John Woo, an applied AI scientist at STScI, on automated mission classification using large language models. The talk begins with the motivation: the rapid growth of astronomical literature makes manual classification infeasible, yet accurate classification is crucial for assessing the scientific impact of missions like HST and JWST. Woo describes a three-stage LLM system: keyword filtering for recall, a reranking stage to prune false positives, and a final classification with structured extraction and reasoning. He emphasizes the importance of evaluation, detailing the creation of a golden sample through human annotation and the use of metrics like precision, recall, and F1. He presents results for identifying JWST science papers, which is used to enforce DOI compliance. The talk concludes with a discussion of the system’s adaptability and the need to balance automation with human oversight.

149 words

Critical Evaluation

Value of the Information & Strength of the Argument

The presentation provides valuable insights into a practical application of LLMs for scientific literature classification. The argumentation is solid, grounded in the speaker’s direct experience and a clear understanding of the challenges. The three-stage system design is logical and well-motivated, addressing issues of scalability and accuracy. The emphasis on evaluation and the creation of a golden sample demonstrates a rigorous approach. The talk also highlights the broader importance of tracking scientific output for assessing mission impact, adding value beyond the technical details.

91 words

Title / Content Match

The title accurately reflects the content: a session from a workshop on language AI in space sciences, featuring a talk on automated mission classification.

Quality & Reliability

8/10

The presentation is by a domain expert (applied AI scientist at STScI) and describes a concrete system with evaluation methodology. The content is technical and grounded in practical experience, but lacks peer-reviewed citations and detailed quantitative results in the transcript.

Key Moments

Cited Sources

  • The Value of the Mikulski Archive for Space Telescopes (MAST) — Referenced as a recent paper by Dick Shaw et al. on the value of the MAST archive, showing that 30% of JWST science is archival.

Concurring Sources

Contribution & Novelties

The talk presents a practical, scalable approach to classifying astronomical literature using LLMs, with a focus on evaluation and human-in-the-loop validation. The three-stage system (keyword filtering, reranking, structured classification) is a novel combination that balances recall and precision. The emphasis on creating a golden sample and comparing LLM performance to human annotators provides a robust framework for deployment.

Pour aller plus loin :

  • Astrophysics Data System (ADS) — The bibliographic database used for the literature queries.
  • MAST Archive — The Mikulski Archive for Space Telescopes, central to the discussion of archival science.
  • JWST DOI Policy — The policy requiring JWST papers to reference a MAST DOI, as mentioned in the talk.

111 words

Radar Profile

The radar profile shows high scores in quantity, quality, and technical level, with a slightly lower but still strong reliability score. This indicates a technically dense and informative presentation, but with some limitations in source citation and verification.

Reliability 7/10