arXiv:2602.21435cs.CV2026-02被引 1

让视觉语言模型在理解与生成间动态交替,实现真正协同。

Synergizing Understanding and Generation with Interleaved Analyzing-Drafting Thinking

  • 设计分析-草稿交替循环,让理解与生成互相迭代优化。
  • 在多个基准上提升理解与生成性能,且适配多种模型架构。
  • 适合希望提升多模态模型协同能力的研究者与开发者。

统一视觉语言模型(UVLMs)旨在通过单一框架同时支持理解与生成任务。然而,现有方法多聚焦于架构统一,忽视了任务求解过程中两种能力的显式交互,导致理解与生成被视为并行技能而非协同过程。为实现真正的协同,我们提出一种交错分析-草稿问题求解循环(AD-Loop),通过动态交替文本思考与视觉思考,使模型能迭代优化理解和输出,实现深度协同。为训练该机制,我们设计两阶段策略:先在交错思维数据上进行监督学习以初始化交替行为,再通过强化学习促进自适应与自主控制。大量实验表明,AD-Loop在标准理解与生成基准上持续提升性能,且具备强泛化能力,适用于多种UVLM架构。视觉分析进一步验证了隐式视觉思考的有效性。结果表明,AD-Loop是一种原则性强且广泛适用的协同策略。项目页面见 https://sqwu.top/AD-Loop。

原文摘要 · Abstract (English)

Unified Vision-Language Models (UVLMs) aim to advance multimodal learning by supporting both understanding and generation within a single framework. However, existing approaches largely focus on architectural unification while overlooking the need for explicit interaction between the two capabilities during task solving. As a result, current models treat understanding and generation as parallel skills rather than synergistic processes. To achieve real synergy, we introduce the interleaved Analyzing-Drafting problem-solving loop (AD-Loop), a new think paradigm that dynamically alternates between analytic and drafting operations. By interleaving textual thoughts with visual thoughts, AD-Loop enables models to iteratively refine both comprehension and outputs, fostering genuine synergy. To train this mechanism, we design a two-stage strategy: supervised learning on interleaved thought data to initialize alternation, followed by reinforcement learning to promote adaptive and autonomous control. Extensive experiments demonstrate that AD-Loop consistently improves performance across standard benchmarks for both understanding and generation, with strong transferability to various UVLMs architectures. Visual analyses further validate the effectiveness of implicit visual thoughts. These results highlight AD-Loop as a principled and broadly applicable strategy for synergizing comprehension and creation. The project page is at https://sqwu.top/AD-Loop.

多模态协同推理生成优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。