统一多模态推理范式,让模型自动生成中间图像来完成复杂任务。
Omni-R1: Towards the Unified Generative Paradigm for Multimodal Reasoning
- 通过生成中间图像实现多种多模态推理技能的统一。
- Omni-R1在多个任务上表现优异,Omni-R1-Zero无需标注数据也能达到同等水平。
- 适合研究多模态生成与通用推理的学者,尤其关注零样本场景。
多模态大语言模型在多模态推理方面取得显著进展。早期方法仅依赖文本推理,近期研究虽引入多模态信息,但通常采用单一任务特定的推理模式,限制了跨任务泛化能力。实际上,存在大量需要多样化推理技能的任务,如定位图像特定区域或标记对象。为此,我们提出统一生成式多模态推理范式,通过在推理过程中生成中间图像来融合多样化的推理能力。我们以Omni-R1为例,采用两阶段SFT+RL框架,结合感知对齐损失和感知奖励,实现功能型图像生成。此外,我们提出Omni-R1-Zero,通过从纯文本推理数据中自举逐步可视化,消除对多模态标注的需求。实验表明,Omni-R1实现了广泛多模态任务上的统一生成推理,而Omni-R1-Zero平均性能可媲美甚至超越Omni-R1,揭示了生成式多模态推理的前景。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) are making significant progress in multimodal reasoning. Early approaches focus on pure text-based reasoning. More recent studies have incorporated multimodal information into the reasoning steps; however, they often follow a single task-specific reasoning pattern, which limits their generalizability across various multimodal tasks. In fact, there are numerous multimodal tasks requiring diverse reasoning skills, such as zooming in on a specific region or marking an object within an image. To address this, we propose unified generative multimodal reasoning, which unifies diverse multimodal reasoning skills by generating intermediate images during the reasoning process. We instantiate this paradigm with Omni-R1, a two-stage SFT+RL framework featuring perception alignment loss and perception reward, thereby enabling functional image generation. Additionally, we introduce Omni-R1-Zero, which eliminates the need for multimodal annotations by bootstrapping step-wise visualizations from text-only reasoning data. Empirical results show that Omni-R1 achieves unified generative reasoning across a wide range of multimodal tasks, and Omni-R1-Zero can match or even surpass Omni-R1 on average, suggesting a promising direction for generative multimodal reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。