让图文交替推理统一成决策过程,用强化学习同时优化文本和图像生成。
Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process

- 将多轮图文交互建模为统一马尔可夫决策过程,实现跨模态联合优化。
- 在空间推理与视觉感知任务上优于多个基线,提升关键视觉分支的生成质量。
- 适合研究多模态生成、强化学习与视觉语言模型融合的学者参考。
统一多模态模型(UMMs)展现出令人期待的图文交替推理能力,但通过强化学习(RL)有效优化此类多轮生成仍是未解难题。现有方法仅对文本步骤应用强化学习,图像生成则依赖监督代理,导致策略梯度无法贯穿异构模态的完整推理轨迹,使强化学习在UMMs中的潜力未被充分挖掘。本文提出 extbf{BRAID}(Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process),将多轮文本-图像-文本推理视为统一的马尔可夫决策过程(MDP),通过单一、严谨的强化学习目标实现文本与视觉生成的联合优化。BRAID 计算共享的轨迹级优势,并一致地传递至文本词元与图像去噪路径,各自通过本模态原生的策略梯度机制进行优化。为进一步解决长程信用分配问题,BRAID 引入视觉语言模型(VLM)裁判,对每一轮中间图像的推理效用进行评分,提供密集的回合级反馈以增强关键视觉分支的学习。在空间推理与视觉感知基准上的实验表明,BRAID 持续优于多种基线,验证了统一的MDP形式化与视觉思维引导对于高效多模态推理至关重要。
原文摘要 · Abstract (English)
Unified multi-modal models (UMMs) have shown promising interleaved text-image reasoning capabilities, yet effectively optimizing such multi-turn generation via reinforcement learning (RL) remains an open challenge. Existing approaches apply RL exclusively to text steps, relegating image generation to supervised surrogates, preventing policy gradients from propagating through the full interleaved trajectory across heterogeneous modalities. This leaves the potential of RL for UMMs largely untapped. In the paper, we introduce \textbf{BRAID} (\textbf{B}ridging inte\textbf{R}le\textbf{A}ved mult\textbf{I}-modal reasoning as a unified \textbf{D}ecision process), a simple framework that casts multi-turn text-image-text reasoning as a unified Markov decision process (MDP), enabling joint optimization of textual and visual generation via a single, principled RL objective. BRAID computes a shared trajectory-level advantage and propagates it coherently into both text tokens and image denoising paths, each optimized through its modality-native policy gradient mechanism. To further address long-horizon credit assignment, BRAID employs a vision-language model (VLM) judge that scores each intermediate image on its reasoning utility, supplying dense turn-level feedback to sharpen learning at critical visual branches. Experiments on spatial reasoning and visual perception benchmarks show that BRAID consistently outperforms various baselines, confirming that a unified MDP formulation with vision-thinking guidance is essential for effective multi-modal reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。