测试视觉与语言信息迁移对物理推理模型的影响,发现模型在图像输入下表现显著下降。
SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoning

- 设计四类渐进式图文变体,精细评估模态转移中的推理能力保留
- 模型在图像输入时平均性能下降,视觉变量定位是主要瓶颈
- 盲训练仍可提升性能,提示模型依赖残留文本线索而非有效视觉证据
我们提出SeePhys Pro,一个细粒度的模态迁移基准,研究当关键信息从文本逐步转移到图像时,模型是否保持相同的推理能力。不同于仅评估单一输入形式的标准视觉必要性基准,SeePhys Pro为每个问题提供四种语义对齐的变体,视觉元素逐步增加。评估显示,当前前沿模型远未达到表示不变推理,性能随信息从语言向图表转移而平均下降,视觉变量定位是最关键的瓶颈。受此推理期脆弱性的启发,我们进一步构建了大规模多模态RLVR训练语料,并使用盲训练作为诊断对照,发现即使训练图像全部遮蔽,模型在未遮蔽验证集上仍能提升性能。通过文本删除、图像遮蔽率和格式饱和控制分析表明,这些增益源于残余文本和分布线索,而非有效的视觉证据。结果强调,评估多模态推理不仅需关注最终答案准确率,还需考察模态转移下的鲁棒性,以及改进是否依赖任务关键的视觉证据。
原文摘要 · Abstract (English)
We introduce SeePhys Pro, a fine-grained modality transfer benchmark that studies whether models preserve the same reasoning capability when critical information is progressively transferred from text to image. Unlike standard vision-essential benchmarks that evaluate a single input form, SeePhys Pro features four semantically aligned variants for each problem with progressively increasing visual elements. Our evaluation shows that current frontier models are far from representation-invariant reasoners: performance degrades on average as information moves from language to diagrams, with visual variable grounding as the most critical bottleneck. Motivated by this inference-time fragility, we further develop large training corpora for multimodal RLVR and use blind training as a diagnostic control, finding that RL with all training images masked can still improve performance on unmasked validation sets. To analyze this effect, text-deletion, image-mask-rate, and format-saturation controls suggest that such gains can arise from residual textual and distributional cues rather than valid visual evidence. Our results highlight the need to evaluate multimodal reasoning not only by final-answer accuracy, but also by robustness under modality transfer and by diagnostics that test whether improvements rely on task-critical visual evidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。