arXiv:2605.18714cs.CVcs.AI2026-05

用图像分割提升多模态模型理解与生成的协同能力

Semantic Generative Tuning for Unified Multimodal Models

论文配图:Semantic Generative Tuning for Unified Multimodal Models
图 1 · 摘自论文原文
  • 以图像分割作为生成代理,统一优化视觉理解与生成
  • 在多个基准上同时提升感知准确率和生成布局精度
  • 适合需要强跨模态对齐的多模态系统开发者

统一多模态模型(UMMs)旨在将视觉理解与生成整合于单一架构中。然而,现有训练范式分别通过稀疏文本信号优化理解、通过密集像素目标优化生成,导致表征空间错位,阻碍两者相互促进。本文首次系统研究生成后训练,提出将分层视觉任务作为生成代理以弥合隔离。实验表明,高层语义任务(尤其是图像分割)为最优代理:相比低层任务带来的纹理干扰,分割提供结构语义,显著提升视觉感知与生成布局保真度。基于此,我们提出语义生成调优(SGT),利用分割作为生成代理实现多模态能力对齐与协同。机制分析显示,SGT显著改善特征线性可分性并优化视觉-文本注意力分配。大量实验证明,SGT在主流基准上持续提升多模态理解和生成保真度。代码已开源。

原文摘要 · Abstract (English)

Unified multimodal models (UMMs) strive to consolidate visual understanding and visual generation within a single architecture. However, prevailing training paradigms independently optimize understanding via sparse text signals and generation through dense pixel objectives. Such a decoupled strategy yields misaligned representation spaces, isolating visual understanding from generation and hindering their mutual reinforcement. This work presents the first systematic investigation into generative post-training, where we formulate hierarchical visual tasks as generative proxies to bridge the isolation in UMMs. Our empirical investigation reveals that high-level semantic tasks, particularly image segmentation, serve as optimal proxies. Unlike low-level tasks that distract models with texture details, segmentation provides structural semantics that significantly enhance both vision-centric perception and generative layout fidelity. Building upon these insights, we introduce Semantic Generative Tuning (SGT), a novel paradigm that leverages segmentation as a generative proxy to align and synergize multimodal capabilities. Mechanistic analyses further demonstrate that SGT fundamentally improves feature linear separability and optimizes visual-textual attention allocation pattern. Extensive evaluations show that SGT consistently improves both multimodal comprehension and generative fidelity across mainstream benchmarks. Our code is available on the https://song2yu.github.io/SGT/.

多模态生成调优图像分割对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。