用强化学习让AI真正理解画图与推理的因果关系。
How RL Unlocks the Aha Moment in Geometric Interleaved Reasoning
- 用强化学习强制模型理解绘图与推理的因果关系
- 在几何推理任务上超越纯文本基线,表现显著提升
- 适合研究多模态推理与模型内化能力的学者
解决复杂几何问题需要在画图与逻辑推导之间反复切换。尽管近期多模态大模型在图像生成和绘图方面表现强劲,但我们发现一个反直觉现象:直接对画图-解题数据进行监督微调(SFT),反而导致推理性能大幅下降,甚至低于纯文本基线。我们指出,SFT 的根本局限在于仅实现分布对齐——模型学会模仿绘图的表面形式,却未掌握绘图与推理步骤间的因果依赖。为此,我们提出 Faire(功能性对齐框架),通过引入三项因果约束,推动模型从表层模仿转向功能对齐。大量实验表明,Faire 引发了模型行为的质变:绘图被真正内化,从而在挑战性几何推理基准上取得竞争力表现。
原文摘要 · Abstract (English)
Solving complex geometric problems inherently requires interleaved reasoning: a tight alternation between constructing diagrams and performing logical deductions. Although recent Multimodal Large Language Models (MLLMs) have demonstrated strong capabilities in visual generation and plotting, we identify a counter-intuitive and underexplored phenomenon. Naively applying Supervised Fine-Tuning (SFT) on interleaved plot-solution data leads to a substantial degradation in reasoning performance compared to text-only baselines. We argue that this failure stems from a fundamental limitation of SFT, which primarily induces distributional alignment: the model learns to reproduce the surface format of interleaved plotting but fails to internalize the causal dependency between the generated plot and reasoning steps. To overcome this limitation, we propose Faire (Functional alignment for interleaved reasoning), a reinforcement learning framework that enforces three casual constraints to move beyond superficial imitation toward functional alignment. Extensive experiments show that Faire induces a qualitative shift in model behavior in which the plotting is effectively internalized, yielding competitive performance on challenging geometric reasoning benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。