arXiv:2603.19807cs.CVcs.AI2026-03

用语义对齐监督提升多模态模型生成质量与对齐效果

Enhancing Alignment for Unified Multimodal Models via Semantically-Grounded Supervision

  • 构建视觉定位图,生成互补监督信号
  • 提升生成保真度和跨模态对齐,显著优于基线
  • 适合需要高质量多模态生成的科研与应用

统一多模态模型(UMMs)将多模态理解与生成集成于统一框架中,但现有生成训练范式存在固有局限。本文提出语义对齐监督(SeGroS),一种微调框架,以解决UMMs中的粒度不匹配与监督冗余问题。核心是设计新型视觉定位图,构建两种互补监督信号:其一,通过语义视觉提示弥补文本提示稀疏性;其二,生成语义对齐的损坏输入,仅在核心对齐区域限制重建损失,从而强化掩码型UMMs的监督。在GenEval、DPGBench和CompBench上的广泛评估表明,SeGroS显著提升了多种UMM架构的生成保真度与跨模态对齐能力。

原文摘要 · Abstract (English)

Unified Multimodal Models (UMMs) have emerged as a promising paradigm that integrates multimodal understanding and generation within a unified modeling framework. However, current generative training paradigms suffer from inherent limitations. We present Semantically-Grounded Supervision (SeGroS), a fine-tuning framework designed to resolve the granularity mismatch and supervisory redundancy in UMMs. At its core, we propose a novel visual grounding map to construct two complementary supervision signals. First, we formulate semantic Visual Hints to compensate for the sparsity of text prompts. Second, we generate a semantically-grounded Corrupted Input to explicitly enhance the supervision of masking-based UMMs by restricting the reconstruction loss to core text-aligned regions. Extensive evaluations on GenEval, DPGBench, and CompBench demonstrate that SeGroS significantly improves generation fidelity and cross-modal alignment across various UMM architectures.

多模态生成模型对齐优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。