用循环一致性强化学习提升多模态推理的准确性
R-C2: Cycle-Consistent Reinforcement Learning Improves Multimodal Reasoning
- 通过跨模态循环约束,让模型自我校正内部表示
- 在多个数据集上推理准确率最高提升7.6分
- 适合研究多模态对齐与自洽模型构建的学者
鲁棒的感知与推理需要跨感官模态的一致性。然而当前多模态模型常违背这一原则,对同一概念的视觉和文本表示产生矛盾预测。我们不采用传统投票机制掩盖这些错误(该机制可能放大系统性偏差),而是利用跨模态不一致作为丰富的自然学习信号。提出RC2框架,通过强化学习强制执行跨模态循环一致性:要求模型进行反向推理、切换模态,并通过正向推理可靠重建答案,从而获得密集的、无需标签的奖励信号。该循环约束促使模型自主对齐内部表示。优化此结构可缓解模态特异性错误,在多个基准上推理准确率最高提升7.6点。结果表明,高级推理不仅源于数据规模扩展,更来自对世界结构化一致理解的强制。
原文摘要 · Abstract (English)
Robust perception and reasoning require consistency across sensory modalities. Yet current multimodal models often violate this principle, yielding contradictory predictions for visual and textual representations of the same concept. Rather than masking these failures with standard voting mechanisms, which can amplify systematic biases, we show that cross-modal inconsistency provides a rich and natural signal for learning. We introduce RC2, a reinforcement learning framework that resolves internal conflicts by enforcing cross-modal cycle consistency. By requiring a model to perform backward inference, switch modalities, and reliably reconstruct the answer through forward inference, we obtain a dense, label-free reward. This cyclic constraint encourages the model to align its internal representations autonomously. Optimizing for this structure mitigates modality-specific errors and improves reasoning accuracy by up to 7.6 points. Our results suggest that advanced reasoning emerges not only from scaling data, but also from enforcing a structurally consistent understanding of the world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。