CLJC Session 19 - D-Fusion: Preference Optimization for Visually Consistent Diffusion Models

CLJC Session 19 - D-Fusion: Preference Optimization for Visually Consistent Diffusion Models

🎙 Mobina Poulaei 👥 1K 📅 August 18, 2025 ⏱ 69 min 👁 24 📄 original study 🧭 2026-08-17
Available in: English (current) Français

Keywords

D-FusionDirect Preference Optimizationdiffusion modelstext-to-image alignmentself-attention fusion

Summary

The presentation introduces D-Fusion, a method to improve text-image alignment in diffusion models by addressing the issue of visual inconsistency in preference pairs used for Direct Preference Optimization (DPO). The speaker explains that DPO requires pairs of good and bad images, but due to different initial noises, these pairs often differ in background, lighting, and structure, leading the model to learn spurious correlations. D-Fusion constructs visually consistent pairs by using mask-guided self-attention fusion. The process involves extracting cross-attention masks from the first up-sampling layer to identify object regions, then fusing self-attention maps from a reference image (good alignment) with the base image (poor alignment) to generate a target image that maintains the background of the base but adopts the object details from the reference. This is done for each denoising step, preserving the denoising trajectory. The method is evaluated using CLIP score and FID, showing improvements in alignment across various RL setups. The presentation also includes a discussion on the challenges of online vs. offline RL and the limitations of using CLIP as a reward model.

176 words

Critical Evaluation

Value of the Information & Strength of the Argument

The presentation provides a thorough explanation of the D-Fusion method, highlighting its novelty in addressing visual inconsistency in DPO for diffusion models. The argumentation is solid, with clear reasoning about why existing methods fail and how D-Fusion overcomes these issues. The speaker also engages in critical discussion, questioning the choice of XOR for mask extraction and the reliance on CLIP score, which adds depth to the analysis. However, the presentation is somewhat informal and includes digressions that may distract from the core message.

Scientific Rigor, Source Quality, Title Accuracy

The presentation is based on a peer-reviewed paper accepted at ICML 2025, which adds credibility. The speaker references the paper and its arXiv link, but does not cite additional sources. The title accurately reflects the content. The presentation includes a critical discussion of the method’s limitations, such as the potential for spurious features and the use of CLIP score, which demonstrates scientific rigor. However, the informal style and lack of structured citations may reduce the perceived rigor.

175 words

Title / Content Match

The title accurately reflects the content, which is a detailed presentation of the D-Fusion paper.

Quality & Reliability

7/10

The presentation is based on a peer-reviewed paper (ICML 2025) and provides a detailed technical walkthrough of the method, including algorithmic steps and experimental setup. However, the presentation is informal, with some digressions and critical discussions that are not fully structured. The presenter demonstrates deep understanding but the delivery is somewhat disorganized.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

The presentation offers a detailed walkthrough of the D-Fusion method, which addresses a critical issue in applying DPO to diffusion models: the lack of visual consistency in preference pairs. By constructing consistent pairs via mask-guided self-attention fusion, D-Fusion improves text-image alignment without requiring additional human feedback. The method is scalable and can be integrated into existing RL setups. The presentation also provides insights into the challenges of using CLIP as a reward model and the importance of preserving denoising trajectories.

Pour aller plus loin :

144 words

Radar Profile

The radar profile shows high scores in technical depth and information quantity, reflecting the detailed algorithmic explanation and comprehensive coverage of the method. The quality and reliability scores are moderate, indicating a solid but not flawless presentation, with some informal digressions and critical discussions that may affect clarity.

Reliability 7/10