让数学题的图文推理更精准,按需分配视觉监督信号。
MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning

- 根据每道题的视觉依赖程度动态调整监督强度
- 在数学推理中实现渐进式视觉感知提升,准确率显著提高
- 适合研究多模态推理与细粒度视觉理解的学者
链式思维(CoT)已从纯语言领域扩展到多模态场景,但现有方法常将视觉输入视为同质或辅助信号,未能捕捉文本与图像在数学解题中的复杂且样本特异的依赖关系。这导致两个核心问题:其一,视觉内容的监督信号过于泛化和粗粒度,无法适配每个样本中视觉信息的实际必要性;其二,当视觉奖励被统一应用而未区分输入间的互补关系时,训练反馈变得不准确。这些限制阻碍了模型实现精确的多模态推理。本文提出一种建模数学推理中细粒度视觉依赖的框架。首先构建MathVis-Fine数据集,通过添加视觉依赖评分实现细粒度视觉标注。基于此数据集,引入两阶段渐进式视觉增强训练范式,依据样本内在的视觉依赖水平平衡答案正确性奖励与视觉定位奖励,从而缓解奖励偏差、提升监督精度。大量实验表明,该框架能根据视觉依赖程度逐步增强视觉感知,为多模态数学推理提供更精确的训练机制。数据集将在论文录用后公开。
原文摘要 · Abstract (English)
Chain-of-Thought (CoT) reasoning has extended from purely linguistic domains to multimodal scenarios; however, existing approaches often treat visual inputs as homogeneous or auxiliary signals, failing to capture the intricate and sample-specific dependencies between text and images in mathematical problem-solving. This gives rise to two core issues: first, the supervisory signals for visual content are generalized and coarse-grained, lacking adaptation to the actual necessity of visual information in each sample; second, training feedback becomes inaccurate when visual rewards are uniformly applied without distinguishing the complementary relationships among inputs. These limitations hinder models from achieving precise multimodal reasoning. In this work, we propose a framework for modeling fine-grained visual dependencies in mathematical reasoning. We first construct the MathVis-Fine dataset, augmenting fine-grained visual annotations with visual dependency ratings. Building upon this dataset, we introduce a two-stage progressive visual enhancement training paradigm that balances answer correctness rewards and visual grounding rewards according to the intrinsic visual dependency level of each sample, thereby mitigating reward bias and improving supervision accuracy. Extensive experiments demonstrate that the MathVis-Fine framework effectively enhances visual perception progressively based on visual dependency, offering a more precise training framework for multimodal mathematical reasoning. We will release the dataset upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。