arXiv:2602.14147cs.CV2026-02被引 4

LaViDa-R1用统一框架让多模态扩散语言模型更强推理能力。

LaViDa-R1: Advancing Reasoning for Unified Multimodal Diffusion Language Models

  • 用统一后训练框架融合监督微调与多任务强化学习。
  • 在视觉数学、图像定位等任务上表现优异,支持图像编辑。
  • 适合研究多模态生成与复杂推理的开发者或学者。

扩散语言模型(dLLMs)成为自回归大模型的有力替代。最新研究将其拓展至多模态理解与生成任务。本文提出LaViDa-R1,一种通用多模态推理扩散语言模型。不同于以往通过任务特定强化学习构建推理dLLMs的方法,LaViDa-R1以统一方式整合多种多模态理解与生成任务。其采用新型统一后训练框架,无缝融合监督微调(SFT)与多任务强化学习(RL),并引入答案强制、树搜索和互补似然估计等新训练技术,提升效果与可扩展性。大量实验表明,LaViDa-R1在视觉数学推理、强推理型定位及图像编辑等广泛多模态任务中表现突出。

原文摘要 · Abstract (English)

Diffusion language models (dLLMs) recently emerged as a promising alternative to auto-regressive LLMs. The latest works further extended it to multimodal understanding and generation tasks. In this work, we propose LaViDa-R1, a multimodal, general-purpose reasoning dLLM. Unlike existing works that build reasoning dLLMs through task-specific reinforcement learning, LaViDa-R1 incorporates diverse multimodal understanding and generation tasks in a unified manner. In particular, LaViDa-R1 is built with a novel unified post-training framework that seamlessly integrates supervised finetuning (SFT) and multi-task reinforcement learning (RL). It employs several novel training techniques, including answer-forcing, tree search, and complementary likelihood estimation, to enhance effectiveness and scalability. Extensive experiments demonstrate LaViDa-R1's strong performance on a wide range of multimodal tasks, including visual math reasoning, reason-intensive grounding, and image editing.

多模态扩散模型推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。