CLJC Session 22 - Attention Prompting on Image for Large Vision-Language Models

CLJC Session 22 - Attention Prompting on Image for Large Vision-Language Models

🎙 Robust and Interpretable Machine Learning Lab 👥 1K 📅 November 18, 2025 ⏱ 23 min 👁 14 📄 literature review 🧭 2026-08-16
Available in: English (current) Français

Keywords

Attention PromptingLVLMVisual PromptingCLIPLLaVA

Summary

This session presents the paper ‘Attention Prompting on Image for Large Vision-Language Models’ (arXiv:2409.17143). The speaker explains the motivation: LVLMs often struggle to follow text instructions when interpreting images. The proposed method, Attention Prompting on Image (API), overlays a text-query-guided attention heatmap onto the input image to help the model focus on relevant regions. The speaker details two approaches for generating the attention map: using a separate CLIP model or using the LVLM itself. They discuss the importance of selecting appropriate layers and handling general vs. specific queries. Experimental results show modest improvements on benchmarks like MM-Vet and LLaVA-Wild. The speaker provides critical analysis, noting the gains are marginal and that the method is training-free. They also mention limitations and potential future directions, such as using open-source models and exploring compositional tasks.

132 words

Critical Evaluation

Value of the Information & Strength of the Argument

The presentation provides a clear and detailed explanation of the proposed method, including its motivation and implementation. The speaker effectively communicates the core idea and supports it with examples and experimental results. However, the argumentation is somewhat informal and lacks deep critical evaluation of the method’s limitations. The speaker does mention that the improvements are marginal and that the method may not be optimal for all tasks, but could have delved deeper into potential weaknesses and comparisons with other approaches.

89 words

Title / Content Match

The title accurately reflects the content, which is a session dedicated to the paper on Attention Prompting on Image for LVLMs.

Quality & Reliability

7/10

The presentation is a thorough review of a specific paper, with clear explanations of the method and results. The speaker is knowledgeable and provides critical analysis, but the presentation is informal and lacks rigorous verification of claims.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

The presentation offers a clear and accessible explanation of the API method, highlighting its novelty as a training-free visual prompting technique that leverages attention maps. It provides a balanced view of the method’s strengths and limitations.

Pour aller plus loin :

  • Visual Prompting — General concept of prompting in AI.
  • CLIP — The CLIP model used for generating attention maps.
  • LLaVA — The LVLM used in the paper’s experiments.

69 words

Radar Profile

The radar profile shows high scores in technical level and information quality, but lower in reliability and quantity, reflecting the informal nature of the presentation and the limited scope of sources.

Reliability 6/10