LLM-powered exploratory text analysis at scale

LLM-powered exploratory text analysis at scale

🎙 Mian Zhong 👥 4K 📅 March 30, 2026 ⏱ 34 min 👁 34 📄 original study 🧭 2026-08-16
Available in: English (current) Français

Keywords

HiCodeinductive codingtopic modelingdata selectionLLM

Summary

Mian Zhong presents her research on scaling inductive coding for content analysis using large language models (LLMs). She introduces HiCode, a two-part pipeline that generates labels from text and then clusters them into themes. The motivation is to handle large text corpora, such as litigation documents, where manual analysis is infeasible. HiCode is evaluated against baselines like TopicGPT and LDA, showing improved performance on theme-level and second-level metrics. A case study on opioid sales strategies demonstrates its practical utility. The second part of the talk addresses data selection strategies, comparing keyword search, random sampling, and semantic retrieval methods. Results show that semantic retrieval yields more relevant topics but may reduce diversity. The presentation concludes with future directions, including model specialization, better evaluation metrics, and user-centric tool design.

127 words

Critical Evaluation

Value of the Information & Strength of the Argument

The presentation provides valuable insights into the application of LLMs for text analysis, introducing a novel pipeline (HiCode) and systematic evaluation of data selection strategies. The argumentation is solid, supported by quantitative results and a case study. However, the presentation is concise and lacks deep discussion of limitations and potential biases.

Scientific Rigor, Source Quality, Title Accuracy

The research appears methodologically sound, with clear definitions and evaluation metrics. The sources cited are not explicitly mentioned in the talk, but the description references the Opioid Industry Documents Archive and the Track COVID dataset. The title accurately reflects the content, focusing on LLM-powered text analysis at scale.

114 words

Title / Content Match

The title accurately reflects the content, focusing on LLM-powered text analysis at scale.

Quality & Reliability

7/10

The presentation is based on original research with a clear methodology, but lacks detailed peer-reviewed publication and full transparency on evaluation metrics.

Key Moments

Cited Sources

Concurring Sources

  • TopicGPT — Used as a baseline in the study.

Contribution & Novelties

The presentation introduces HiCode, a novel LLM-based pipeline for inductive coding at scale, and systematically evaluates data selection strategies for downstream text analysis. It highlights the trade-off between relevance and diversity in topic modeling. For further exploration, consider the following:

69 words

Radar Profile

The radar profile shows high scores in quantity of information and technical level, indicating a dense and technical presentation. Quality of information and reliability are slightly lower, reflecting the lack of detailed methodological exposition and peer-reviewed publication.

Reliability 7/10