通过锚点注意力提升多模态模型推理的视觉准确性
Credit Where It is Due: Cross-Modality Connectivity Drives Precise Reinforcement Learning for MLLM Reasoning
- 识别出仅15%高连接度令牌作为视觉锚点,主导推理过程
- 新方法使32B模型在MathVista上达80.2分,超72B基线
- 适合关注多模态强化学习与视觉-语言对齐的研究者
基于可验证奖励的强化学习(RLVR)显著提升了多模态大语言模型(MLLMs)的推理能力,但视觉证据如何融入推理过程仍不明确。本文从跨模态注意力连通性角度研究多模态RLVR,发现仅有约15%的标记表现出强视觉-文本耦合。这些高连通性标记作为锚点,将推理锚定在图像上,其余多数遵循语言模式。在RLVR训练中,信用分配自然集中于这些锚点,使其视觉锚定能力随时间增强。基于此,我们提出轻量级框架AT-RL,通过注意力拓扑的图聚类选择性强化高连通性标记。在系列模型(3B-32B)上评估显示,AT-RL仅引入1.2%开销,即让32B模型在MathVista上达到80.2分,超越72B-Instruct基线,且在STEM、视频和通用任务中均获一致提升。相反,仅训练低连通性标记会导致严重退化,证实有效多模态强化学习依赖于对视觉锚点的精准信用分配。本工作表明,推理质量取决于跨模态锚定的保真度,而非标记数量。
原文摘要 · Abstract (English)
Reinforcement Learning with Verifiable Rewards (RLVR) has significantly advanced the reasoning capabilities of Multimodal Large Language Models (MLLMs), yet how visual evidence is integrated during reasoning remains poorly understood. We explore multimodal RLVR through the lens of cross-modal attention connectivity and find that only a small fraction of tokens (approximately 15%) exhibit strong visual-textual coupling. These high-connectivity tokens act as anchors that ground reasoning in the image, while the majority follow linguistic patterns. During RLVR training, credit assignment naturally concentrates on these anchors, sharpening their visual grounding over time. Building on this insight, we propose Anchor-Token Reinforcement Learning (AT-RL), a lightweight framework that selectively reinforces high-connectivity tokens via graph-based clustering of attention topology. Evaluated across the series (3B-32B), AT-RL introduces only 1.2% overhead yet enables the 32B model to surpass the 72B-Instruct baseline on MathVista (80.2), with consistent gains observed across STEM, video and general tasks. Conversely, training solely on low-connectivity tokens causes severe degradation, confirming that effective multimodal RL hinges on precise credit assignment to visual anchors. Our work reveals that reasoning quality is governed not by token quantity but by the fidelity of cross-modal anchoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。