arXiv:2605.25437cs.CV2026-05被引 1

多源视觉推理中,新方法通过锚定单源表现提升融合效果。

Does Seeing More Mean Knowing More? Mono-Anchored Advantage Normalization for Multi-Source Visual Reasoning

论文配图:Does Seeing More Mean Knowing More? Mono-Anchored Advantage Normalization for Multi-Source Visual Reasoning
图 1 · 摘自论文原文
  • 以单源奖励为动态锚点,量化多源融合的信息增益。
  • 在多个数据集上分别提升3.2%和4.9%性能。
  • 适合处理红外、深度等异质多源输入场景。

基于可验证奖励的强化学习(RLVR)在视觉推理任务中取得了显著进展。然而,在处理多源输入时,现有方法往往将信息简单叠加,缺乏明确机制来判断多源融合是否带来信息增益或引入干扰。当各源在物理特性与语义上差异较大(如红外与深度图像)时,其动态交互建模能力不足,导致性能反而劣于单源推理,尤其在某一源占主导信号时更为明显。为此,我们提出MARS框架,将每种视觉模态视为独立信息源,通过将单源奖励作为动态锚点,显式地将多源融合带来的信息增益融入优势归一化,并自适应增强源间协同、抑制潜在噪声或冲突。理论分析表明,该方法能有效量化梯度估计中的信息增益,实现一致的模态调节。实证结果表明,在GRPO和DAPO基准上,跨多样数据集分别取得3.2%和4.9%的性能提升,验证了方法的有效性。

原文摘要 · Abstract (English)

Visual reasoning through reinforcement learning with verifiable rewards (RLVR) has achieved remarkable progress. However, when dealing with multi-source inputs, existing approaches tend to treat them as a mere accumulation of information, lacking explicit mechanisms to distinguish whether integrating additional sources yields information gain or introduces interference. Therefore, they struggle to effectively model dynamic interaction when integrating multiple sources, particularly when they differ significantly in physical properties and semantics, e.g., infrared and depth, leading to inferior performance to mono-source reasoning when a certain source holds the dominant signal. To address this issue, we propose MARS, a novel mono-anchored multi-source reasoning framework that models each visual modality as an independent information source. Specifically, by treating mono-source rewards as dynamic anchors, our method explicitly incorporates the information gain introduced by multi-source fusion into advantage normalization and adaptively emphasizes mutual promotion between sources while suppressing potential noise or conflicts during RLVR. From theoretical analysis, our method effectively quantifies information gain introduced by multi-source integration in gradient estimation, enabling consistent modality regulation. Empirical results also show impressive 3.2% and 4.9% performance gains on GRPO and DAPO across diverse datasets, confirming effectiveness of our method.

视觉推理多源融合强化学习优势归一化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。