arXiv:2606.12886cs.CVcs.AI2026-06

通过分步强化学习,解决多模态推理中图文交替时信息脱节问题。

Bridging Modal Isolation in Interleaved Thinking: Supervising Modality Transitions via Stepwise Reinforcement

论文配图:Bridging Modal Isolation in Interleaved Thinking: Supervising Modality Transitions via Stepwise Reinforcement
图 1 · 摘自论文原文
  • 将推理过程拆解为原子操作,量化模态转换时的错误损失
  • 在4个视觉谜题数据集上,跨模态一致性与任务准确率显著提升
  • 适合需要精准图文交互的复杂推理场景研究者

交错推理中,统一的多模态模型在文本推理与视觉生成间交替,已在空间与物理任务中展现潜力。但在复杂长链场景下,我们发现根本性缺陷:生成图像偏离文本语境,后续文本又忽略视觉证据,导致两模态交替却无法真正互推。我们称此为‘模态隔离’,归因于模态边界处的信息持续丢失。我们将每个推理周期分解为原子操作,定义模态转换损失,量化跨模态幻觉(文本到图像)与视觉利用不足(图像到文本)在每处边界的程度。提出MoTiF(模态转换保真度)框架,分两阶段训练:反射式SFT使模型能检测并恢复错误的视觉输出;Flow-GRPO通过强化学习提升图像生成保真度。所有训练信号均来自转换层级保真度,而非最终任务精度。在四个视觉谜题基准上,这种转换级监督显著提升跨模态一致性和最终任务准确率。结果表明,有效交错推理需在模态边界进行显式结构化监督,而非仅靠模型规模或端到端优化。

原文摘要 · Abstract (English)

Interleaved thinking, where a unified multimodal model alternates between textual reasoning and visual generation, has shown promise on spatial and physical tasks. However, in complex long-chain scenarios, we identify a fundamental failure mode: generated images diverge from the textual context while subsequent text ignores the visual evidence, causing the two modalities to alternate without genuinely informing each other. We term this Modal Isolation and attribute it to compounding information loss at modality boundaries. We decompose each reasoning cycle into atomic operations and define modality transition loss, quantifying cross-modal hallucination (text-to-image) and visual utilization deficit (image-to-text) at each boundary. We propose MoTiF (Modality Tiransition Fidelity), a two-stage training framework that directly optimizes these transitions: Reflective SFT trains the model to detect and recover from erroneous visual outputs; Flow-GRPO improves image generation fidelity via reinforcement learning. All training signals in MoTiF derive from transition-level fidelity rather than end-task accuracy. Across four visual puzzle benchmarks, this transition-level supervision substantially improves both cross-modal coherence and final task accuracy. The results demonstrate that effective interleaved reasoning requires explicit structural supervision at modality boundaries, not merely scaling or end-task optimization.

多模态推理强化学习模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。