arXiv:2605.05781cs.CVcs.AI2026-05被引 2

让理解任务指导生成,提升多模态模型的协同效果。

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision

论文配图:Steering Visual Generation in Unified Multimodal Models with Understanding Supervision
图 1 · 摘自论文原文
  • 用理解任务作为监督信号,引导生成过程
  • 在图像生成与编辑中显著提升质量
  • 轻量级设计,适合实际部署

统一的多模态模型旨在弥合理解与生成之间的差距。然而,当前顶尖模型大多采用解耦的理解与生成组件,虽在单任务上表现良好,但削弱了二者相互促进所需的连接,导致潜在协同效应尚不明确。本文提出理解导向的后训练框架(UNO),将理解不仅视为独立任务,更作为直接监督信号来引导生成表示。通过引入编码语义抽象(描述生成)和结构细节(视觉回归)的目标,实现理解到生成的有效梯度传播。大量实验表明,理解任务可有效催化生成性能,显著提升图像生成与编辑效果。

原文摘要 · Abstract (English)

Unified multimodal models are envisioned to bridge the gap between understanding and generation. Yet, to achieve competitive performance, state-of-the-art models adopt largely decoupled understanding and generation components. This design, while effective for individual tasks, weakens the connection required for mutual enhancement, leaving the potential synergy empirically uncertain. We propose to explicitly restore this synergy by introducing Understanding-Oriented Post-Training (UNO), a lightweight framework that treats understanding not only as a distinct task, but also a direct supervisory signal to steer generative representations. By incorporating objectives that encode semantic abstraction (captioning) and structural details (visual regression), we enable effective gradient flow from understanding to generation. Extensive experiments on image generation and editing demonstrate that understanding can serve as an effective catalyst for generation.

多模态生成理解引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。