In-distribution AI-generated literature for cultural simulation

In-distribution AI-generated literature for cultural simulation

🎙 Matthew Wilkens 👥 4K 📅 April 7, 2026 ⏱ 74 min 👁 66 📄 expert opinion 🧭 2026-08-16
Available in: English (current) Français

Keywords

cultural simulationliterary generationlarge language modelscomputational humanitiesexperimental history

Summary

Matthew Wilkens, Associate Professor of Information Science at Cornell, presents a vision for experimental humanities using AI to simulate cultural and historical processes. He argues that current humanities research lacks experimental methods due to the impossibility of intervening on historical events. To address this, he proposes building AI systems that are culturally and temporally constrained, validated against historical records, and capable of generating in-distribution literary texts. He demonstrates a proof-of-concept using the ConLit corpus of ~2,800 contemporary novels, employing a promptable document embedding model (Nomic Embed 8B) to map genre distinctions. He shows that document embeddings effectively separate genres and that initial chunks of novels are good proxies for full texts. For generation, he uses GPT-5 with a system instruction and author biography context to produce novels that are more varied and in-distribution compared to prior work. He presents results showing that AI-generated novels, when embedded, fall within the high-status literary region of the UMAP plot, suggesting they are in-distribution. He concludes by outlining ongoing work on full-scale simulations of literary evolution, including the GPT-1914 model trained on pre-1914 texts, and discusses challenges of removing anachronistic knowledge and validating historical reasoning.

191 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into the potential of AI for cultural simulation, offering a concrete methodology and preliminary results. The argumentation is well-structured, moving from a grand challenge to a research vision and then to a proof-of-concept. The speaker acknowledges limitations and prior work, and the use of document embeddings and genre analysis is methodologically sound. However, the results are preliminary and not yet peer-reviewed, and the speaker does not provide detailed quantitative metrics for the in-distribution claim beyond visual inspection of UMAP plots.

Scientific Rigor, Source Quality, Title Accuracy

The talk is scientifically rigorous, referencing prior work by Melanie Walsh, Chakrabarty, and Andrew Piper’s ConLit corpus. The speaker is transparent about the methods and limitations. The title accurately reflects the content, focusing on in-distribution AI-generated literature for cultural simulation. The talk does not include a public advertising segment.

149 words

Title / Content Match

The title accurately reflects the content, focusing on generating in-distribution literary texts for cultural simulation.

Quality & Reliability

8/10

The talk presents a clear research agenda and preliminary results from a credible academic lab, but lacks full methodological details and peer-reviewed validation for the specific claims.

Key Moments

Cited Sources

  • ConLit corpus — A corpus of ~2,800 contemporary novels assembled by Andrew Piper, used as the test corpus for genre analysis.
  • Walsh et al. on poetry generation — Prior work showing limitations of AI in poetry generation, cited as evidence of narrow and shallow outputs.
  • Chakrabarty et al. on human discrimination — Study showing humans can easily distinguish AI-authored literary texts, cited as prior work.
  • GPT-1914 — A model trained on pre-1914 documents from HathiTrust, mentioned as an approach to avoid anachronism.

Concurring Sources

  • ConLit corpus — The corpus is used as a benchmark for genre analysis, and the results align with expectations about genre compactness.
  • Prior work on AI-generated literature — The talk's findings that AI-generated novels are in-distribution contrast with prior work showing narrow outputs, but the speaker builds on that work.

Dissenting Sources

  • Prior work on AI-generated literature — The speaker's claim that AI can generate in-distribution literary texts contradicts the emerging consensus that AI-generated literature is poor.

Contribution & Novelties

The talk presents a novel approach to generating in-distribution literary texts using LLMs, with a focus on cultural simulation. The use of document embeddings to validate genre distinctions and the finding that initial chunks are good proxies are valuable contributions. The ongoing work on full-scale simulations and the GPT-1914 model are also innovative.

Pour aller plus loin :

  • Cultural analytics — A journal on computational methods for cultural research, relevant to the methodological approach.
  • Digital humanities — Overview of the field, providing context for the research.
  • Large language models — Background on the technology used.
  • Experimental history — Concept related to the grand challenge discussed.

105 words

Radar Profile

The radar profile shows high scores in information quantity, quality, and reliability, with a slightly lower technical level, reflecting the talk's balance of conceptual vision and practical methods.

Reliability 8/10

💬 No comments were provided for analysis.