arXiv:2505.18503cs.CV2025-05ACL被引 6

通过自动对齐注意力提升医疗多模态模型的准确性

Focus on What Matters: Enhancing Medical Vision-Language Models with Automatic Attention Alignment Tuning

  • 用SAM和BioMedCLIP生成弱标签,精准定位视觉关键区域
  • 仅微调关键注意力头,使模型在医学问答上准确率提升5.2%
  • 适合医疗AI研发者,尤其关注模型可解释性的团队

医疗大视觉语言模型常在视觉输入上出现注意力分布不佳,导致生成错误或幻觉。现有方法多依赖推理时干预,适应性差且需额外标注。为此,我们提出A$^3$Tune,一种自动注意力对齐微调框架。该方法利用SAM的零样本弱标签,经BioMedCLIP优化为提示感知标签,并选择性调整视觉关键注意力头以增强对齐,同时减少干扰。此外,引入A$^3$MoE模块,实现跨不同提示与图像的自适应参数选择。在医学VQA与报告生成基准上的实验表明,A$^3$Tune优于当前最优基线,显著改善注意力分布与模型性能。

原文摘要 · Abstract (English)

Medical Large Vision-Language Models (Med-LVLMs) often exhibit suboptimal attention distribution on visual inputs, leading to hallucinated or inaccurate outputs. Existing mitigation methods primarily rely on inference-time interventions, which are limited in attention adaptation or require additional supervision. To address this, we propose A$^3$Tune, a novel fine-tuning framework for Automatic Attention Alignment Tuning. A$^3$Tune leverages zero-shot weak labels from SAM, refines them into prompt-aware labels using BioMedCLIP, and then selectively modifies visually-critical attention heads to improve alignment while minimizing interference. Additionally, we introduce a A$^3$MoE module, enabling adaptive parameter selection for attention tuning across diverse prompts and images. Extensive experiments on medical VQA and report generation benchmarks show that A$^3$Tune outperforms state-of-the-art baselines, achieving enhanced attention distributions and performance in Med-LVLMs.

医疗AI注意力对齐多模态微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。