让视觉语言模型在多模态几何题中表现更一致,通过互相学习提升推理能力。
MIRROR: Learning from the Other View for Multi-Modal Reasoning

- 用不同模态视角互为教师,通过反KL损失引导模型优化
- 在几何推理任务上准确率显著提升,跨模态表现更稳定
- 适合研究多模态推理不一致性及模型鲁棒性提升的学者
与具备强推理能力的大语言模型不同,视觉语言模型在视觉推理任务中表现不佳,即使面对文本、图像和图文结合三种等价视图的几何问题。我们发现这些视图常引发不同行为:模型可能从文本解题却在对应图像上失败,或反之。这种不一致性表明不同视图揭示了互补的推理路径与失败模式,而标准多模态后训练未能充分挖掘。为此,我们构建了ODA-Data,一个高质量的成对多模态几何数据集,包含文本主导、图像主导及图文结合视图,配有训练与评估划分,用于分析模态依赖的推理行为。我们提出模态感知的互训优化方法MIRROR,基于强化学习实现自监督。对每个问题,MIRROR评估所有视图,选取表现最佳视图为教师,其余视图以反KL目标向其学习。在多个几何推理基准测试中,MIRROR优于标准强化学习,在跨模态间实现更高准确率与更强一致性。
原文摘要 · Abstract (English)
Unlike large language models (LLMs) that exhibit strong reasoning capabilities, vision-language models (VLMs) struggle with visual reasoning, even on geometry problems that admit equivalent text, diagram, and combined diagram+text views. We show that these views often elicit different behaviors: a model may solve a problem from text but fail on the corresponding diagram, or succeed visually while failing textually. This inconsistency suggests that different views expose complementary reasoning paths and failure modes that standard multimodal post-training does not fully exploit. To study and exploit this phenomenon, we construct ODA-Data, a high-quality paired multimodal geometry dataset with text-dominant, image-dominant, and combined image+text views of the same problems, together with splits for training and evaluating modality-dependent reasoning behaviors. We then develop Modality-Informed Reciprocal Reasoning Optimization (MIRROR), a reinforcement learning approach for improving multimodal reasoning via self supervision. For each problem, MIRROR evaluates the model under all views, selects the best-performing view as a teacher, and trains other views with a reverse-KL objective towards the teacher. Across reasoning benchmarks that evaluate on geometry problems, MIRROR improves over standard RL and yields more accurate and consistent behavior across modalities
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。