通过视觉注意力对齐提升视觉语言模型推理的可解释性与可信度。
Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward
- 用显著图定位图像关键区域,追踪视觉信息在推理中的流动。
- 以显著图与人工标注框重合度为奖励,优化模型关注真实视觉证据。
- 提升模型推理可信度,适合需要可解释性的应用场景。
视觉语言模型在各类任务中表现卓越,但其可信度仍受质疑,常过度依赖文本线索而忽视视觉证据,易生成无根据或虚构的回答。为此,我们提出Saliency-R1框架,增强模型推理的可解释性与忠实性。该方法引入一种新型显著图技术,高效识别生成答案所依赖的关键图像区域,且不增加额外计算开销。该技术可进一步追踪视觉信息如何贯穿推理过程并影响最终输出,揭示思维过程与视觉上下文的一致性。我们采用显著图与人工标注边界框的重叠率作为奖励函数,结合组相对策略优化(GRPO),引导模型在推理时聚焦于相关视觉区域。实验表明,Saliency-R1显著提升了推理的忠实性、可解释性及整体任务性能。
原文摘要 · Abstract (English)
Vision-language models (VLMs) have achieved remarkable success across diverse tasks. However, concerns about their trustworthiness persist, particularly regarding tendencies to lean more on textual cues than visual evidence and the risk of producing ungrounded or fabricated responses. To address these issues, we propose Saliency-R1, a framework for improving the interpretability and faithfulness of VLMs reasoning. Specifically, we introduce a novel saliency map technique that efficiently highlights critical image regions contributing to generated tokens without additional computational overhead. This can further be extended to trace how visual information flows through the reasoning process to the final answers, revealing the alignment between the thinking process and the visual context. We use the overlap between the saliency maps and human-annotated bounding boxes as the reward function, and apply Group Relative Policy Optimization (GRPO) to align the salient parts and critical regions, encouraging models to focus on relevant areas when conduct reasoning. Experiments show Saliency-R1 improves reasoning faithfulness, interpretability, and overall task performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。