CLJC Session 20 - Minority-Focused Text-to-Image Generation via Prompt Optimization

CLJC Session 20 - Minority-Focused Text-to-Image Generation via Prompt Optimization

🎙 Amir Kasaei 👥 1K 📅 September 17, 2025 ⏱ 56 min 👁 45 📄 literature review 🧭 2026-08-16
Available in: English (current) Français

Keywords

minority instancesprompt optimizationdiffusion modelstext-to-imagelikelihood objective

Summary

This journal club session presents a detailed review of the paper ‘Minority-Focused Text-to-Image Generation via Prompt Optimization’ by Amir Kasaei. The presenter explains the problem of bias in text-to-image diffusion models, which tend to generate samples from high-density regions of the data distribution, underrepresenting minority instances. The paper proposes a prompt optimization framework that maintains semantic content while encouraging the emergence of underrepresented visual features. The method introduces a learnable token appended to the prompt, optimized via a likelihood-based objective to guide generation toward minority samples. The presenter discusses the technical details, including the formulation, the role of classifier-free guidance, and the optimization algorithm. Critical analysis highlights potential issues such as local minima and the dependence on initialization. Experimental results show improvements in quality and diversity across multiple diffusion models. The session includes interactive discussion and comparisons with related work.

140 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video provides a thorough and critical analysis of the paper, explaining the motivation, method, and results. The presenter evaluates the strengths and weaknesses, discussing theoretical issues like the local minimum problem and the role of initialization. The argumentation is solid, with clear reasoning and references to the paper’s claims. The discussion adds value by questioning the method’s assumptions and suggesting potential improvements.

Scientific Rigor, Source Quality, Title Accuracy

The presentation is rigorous, with a clear explanation of the method and its theoretical foundations. The sources cited include the paper itself and the presenter’s personal website. The title accurately reflects the content. The discussion is well-structured and demonstrates a deep understanding of the topic. No external sources are mentioned beyond the paper, but the critical analysis adds credibility.

137 words

Title / Content Match

The title accurately reflects the content, which focuses on minority-focused text-to-image generation via prompt optimization.

Quality & Reliability

7/10

The presentation is a detailed technical review of a CVPR 2025 paper, with critical discussion of the method's limitations and theoretical issues. The presenter demonstrates deep understanding and provides critical analysis, but the video is a journal club session with limited external validation.

Key Moments

Cited Sources

Concurring Sources

Dissenting Sources

  • No discordant sources found — The video does not mention any sources that contradict the paper's findings.

Contribution & Novelties

The video provides a critical and detailed review of a recent paper, offering insights into the method’s strengths and weaknesses. The presenter’s analysis of the local minimum problem and initialization sensitivity adds value beyond the paper itself. The discussion also connects the method to broader concepts in prompt optimization and generative models.

Pour aller plus loin :

99 words

Radar Profile

The radar profile shows high scores in quantity and quality of information, with a strong technical level. The reliability is slightly lower, reflecting the critical and somewhat speculative nature of the discussion. Overall, the video is a solid technical review.

Reliability 7/10