arXiv:2502.01419cs.CVcs.AI2025-02ICML被引 22

提升图像描述生成的精准与完整,不增加计算开销

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models

  • 仅增强关键视觉标记的注意力,避免全量放大
  • 在生成长文本时保持视觉注意力清晰,精度与召回率双升
  • 无需训练,适合现有大模型快速优化

详细图像描述对数据生成和视障人士辅助至关重要。高质量描述需兼顾精确性与完整性,当前多模态大语言模型仍面临挑战。本文假设问题源于生成过程中视觉注意力逐渐减弱且变得嘈杂。为此提出 SPARC(选择性渐进注意力重校准)方法,一种无需训练的优化策略,通过有选择地增强关键视觉标记来提升其贡献。基于三大观察:(1)放大所有视觉标记会降低召回率,故仅选择性增强;(2)随生成长度增加,注意力变噪,利用时间步间注意力差异识别关键标记;(3)视觉注意力逐步减弱,需主动强化以维持影响。实验表明,现有方法提升精度却牺牲召回率,而本方法在极低计算开销下同时提升两者。

原文摘要 · Abstract (English)

Detailed image captioning is essential for tasks like data generation and aiding visually impaired individuals. High-quality captions require a balance between precision and recall, which remains challenging for current multimodal large language models (MLLMs). In this work, we hypothesize that this limitation stems from weakening and increasingly noisy visual attention as responses lengthen. To address this issue, we propose SPARC (Selective Progressive Attention ReCalibration), a training-free method that enhances the contribution of visual tokens during decoding. SPARC is founded on three key observations: (1) increasing the influence of all visual tokens reduces recall; thus, SPARC selectively amplifies visual tokens; (2) as captions lengthen, visual attention becomes noisier, so SPARC identifies critical visual tokens by leveraging attention differences across time steps; (3) as visual attention gradually weakens, SPARC reinforces it to preserve its influence. Our experiments, incorporating both automated and human evaluations, demonstrate that existing methods improve the precision of MLLMs at the cost of recall. In contrast, our proposed method enhances both precision and recall with minimal computational overhead.

图像描述注意力机制多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。