让统一多模态模型生成更准,靠的是自我反思激活内在知识。
Free Lunch for Unified Multimodal Models: Enhancing Generation via Reflective Rectification with Inherent Understanding

- 用自我反思链式思维,激活模型理解能力来改进生成。
- 在扩散去噪过程中对齐中间结果与指令理解,提升生成质量。
- 无需训练即可增强现有模型,适合追求高质量生成的用户。
统一多模态模型(UMMs)旨在将视觉理解和生成整合于单一结构中。然而,这类模型存在显著的能力失衡:其理解能力远强于生成能力。这表明模型丰富的内部知识在生成时未被充分激活。受人类‘边画边思考’模式启发,本文提出UniRect-CoT——一种无需训练的统一修正链式思维框架。将UMM中的扩散去噪过程视为内在视觉推理过程,通过将中间结果与模型理解的目标指令对齐,生成自监督信号以修正生成过程。大量实验表明,UniRect-CoT可无缝集成至现有UMMs,在多样复杂任务中显著提升生成质量。
原文摘要 · Abstract (English)
Unified Multimodal Models (UMMs) aim to integrate visual understanding and generation within a single structure. However, these models exhibit a notable capability mismatch, where their understanding capability significantly outperforms their generation. This mismatch indicates that the model's rich internal knowledge, while effective for understanding tasks, remains underactivated during generation. To address this, we draw inspiration from the human ``Thinking-While-Drawing'' paradigm, where humans continuously reflect to activate their knowledge and rectify intermediate results. In this paper, we propose UniRect-CoT, a training-free unified rectification chain-of-thought framework. Our approach unlocks the ``free lunch'' hidden in the UMM's powerful inherent understanding to continuously reflect, activating its internal knowledge and rectifying intermediate results during generation.We regard the diffusion denoising process in UMMs as an intrinsic visual reasoning process and align the intermediate results with the target instruction understood by the model, serving as a self-supervisory signal to rectify UMM generation.Extensive experiments demonstrate that UniRect-CoT can be easily integrated into existing UMMs, significantly enhancing generation quality across diverse complex tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。