让图文理解与生成共享同一目标,提升模型一致性。
STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs

- 通过共享目标状态连接理解和生成路径
- 在图像编辑任务中显著提升描述与图像的一致性
- 适合需要精准图文对齐的研究与应用
统一多模态模型(UMMs)试图在单一架构中整合视觉理解与生成能力,但仅靠结构统一无法保证语义一致。模型可能正确描述目标却生成不一致的图像,暴露了理解与生成间的对齐鸿沟:语言和视觉输出虽处于不同空间,却应遵循相同的目标语义。本文以图像编辑为例,分析给定源图与编辑指令时,模型生成的描述与图像是否共同指向同一目标状态。分析表明,现有UMMs在细粒度实体、属性、空间关系和局部细节上仍存在弱对齐,说明单纯架构统一不足以实现语义融合。为此,提出STBridge框架,通过共享目标状态连接理解与生成:目标描述表达期望的视觉结果,而生成图像则具体实现该结果,取代原有的独立任务路径,构建从目标表达到目标实现的共享信息流。STBridge采用先对齐后优化策略:监督微调建立共享通道,序列强化学习进一步精炼以目标为中心的协调能力。在视觉理解、图像生成与图像编辑多个基准测试中,STBridge均优于初始化模型。对齐分析证实,其有效缩小了描述与生成之间的差距,证明共享目标对齐是提升UMM理解与生成一致性的有效后训练策略。
原文摘要 · Abstract (English)
Unified multimodal models (UMMs) aim to integrate visual understanding and generation within a single architecture, but architectural unification alone does not ensure semantic consistency. A model may describe the intended target correctly while generating an inconsistent edit. This exposes an understanding-generation alignment gap: linguistic and visual outputs live in different spaces, yet should be governed by the same target semantics. We study this gap in image editing, where an instruction defines a target state that can be both described and visually realized. Given a source image and an edit instruction, we compare a UMM's target caption with its edited image to test whether the two outputs converge on the same result. Our analysis shows that existing UMMs remain weakly aligned, especially for fine-grained entities, attributes, spatial relations, and local details, indicating that semantic unification is not achieved by architecture alone. To bridge this gap, we propose STBridge, a shared-target alignment framework that connects understanding and generation through a common target state. Here the target caption expresses the desired visual result, while the edited image realizes it visually, replacing separate task-specific paths with a shared information flow from target expression to target realization. STBridge follows an align-then-optimize strategy: supervised fine-tuning first establishes the shared-target channel, and sequential reinforcement learning further refines target-centered coordination. Across visual understanding, image generation, and image editing benchmarks, STBridge consistently improves over the initialization model. Alignment analysis confirms that STBridge narrows the gap between what the model describes and what it generates, demonstrating shared-target alignment as an effective post-training strategy for bridging understanding and generation in UMMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。