Conférence plénière : Jean Philippe Magué (ENS Lyon, France -- 26 septembre 2025)

Conférence plénière : Jean Philippe Magué (ENS Lyon, France -- 26 septembre 2025)

Humanities, Social Sciences & Thought Mathematics PBMathematics
🎙 Jean Philippe Magué 👥 63 📅 December 2, 2025 ⏱ 41 min 👁 53 📄 expert opinion 🧭 2026-08-16
Available in: English (current) Français

Keywords

genderLLMsociolinguisticsTwitterstereotypes

Summary

Jean Philippe Magué presents a study on whether large language models (LLMs) can infer the gender of Twitter users from their tweets. He begins with a corpus of 650 million tweets from 3 million users, from which he extracted a subset of 4,000 users with manually annotated gender and age. He first tests humans: 180 participants correctly identify the gender of tweet authors 63% of the time, with men being less accurate at identifying female authors. Then he tests GPT-4 on 1,000 tweets (500 from each gender) under three conditions: prompting the model to adopt a male, female, or unknown gender. GPT-4 achieves similar accuracy to humans (63-65%). He analyzes the justifications generated by the model using embeddings and linear discriminant analysis, revealing that the model’s justifications cluster by the gender it was assigned, and that it draws on gender stereotypes. For instance, when the model is female, it describes female-authored tweets as emotionally expressive, while when male, it describes them as ‘unfiltered’. The study highlights how LLMs reproduce and potentially amplify gender stereotypes, and raises questions about the ethical implications of such inference.

184 words

Critical Evaluation

Value of the Information & Strength of the Argument

The value of the information is high: the study uses a large, real-world dataset and a rigorous methodology, including manual annotation and statistical analysis. The argumentation is solid, as the speaker systematically compares human and LLM performance and uses embedding analysis to uncover patterns. However, the presentation is more of a research talk than a fully peer-reviewed study, and some methodological details (e.g., the exact prompt, the selection of extreme tweets) are not fully elaborated. The speaker acknowledges limitations, such as the binary gender approach and the non-representative sample of human participants.

Scientific Rigor, Source Quality, Title Accuracy

The scientific rigor is good: the study is based on a large corpus and uses established methods like LDA and PCA. The sources are not explicitly cited in the talk, but the speaker mentions using the INSEE data and the Voyage AI embedding model. The title is accurate but generic; it does not hint at the specific content. The talk does not include a formal literature review, but it references known sociolinguistic findings. The adequacy between title and content is good, though the title could be more descriptive.

195 words

Title / Content Match

The title accurately describes the content: a plenary conference by Jean Philippe Magué, though it lacks specificity about the topic.

Quality & Reliability

7/10

The presentation is based on a large-scale corpus study with manual annotation and rigorous statistical methods, but the results are preliminary and rely on a binary gender perspective. The speaker acknowledges limitations and ethical concerns, but the methodology is not fully detailed in the talk.

Key Moments

Cited Sources

  • Voyage AI — Embedding model used to vectorize justifications.
  • INSEE — French national statistics institute providing socio-demographic data.

Concurring Sources

  • Voyage AI — Embedding model used in the analysis.

Contribution & Novelties

This talk provides a novel empirical investigation into the ability of LLMs to infer gender from text, comparing human performance and analyzing the stereotypes embedded in the model’s justifications. The use of embedding analysis to reveal how the model’s assigned gender influences its reasoning is innovative. The findings highlight the risk of LLMs perpetuating gender stereotypes.

Pour aller plus loin :

94 words

Radar Profile

The radar profile shows high scores in quantity and quality of information, moderate technical level, and good reliability. This indicates a well-informed presentation with solid data, though not extremely technical or exhaustive in methodology.

Reliability 7/10

💬 No comments were provided for analysis.