arXiv:2509.23760cs.CV2025-09AAAI被引 2

统一视觉语言生成框架,提升复杂指令下的跨模态一致性。

UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and Perception

  • 单扩散模型内融合双流训练,实现视觉与文本语义对齐。
  • 在复杂指令下跨模态一致性显著优于现有基线。
  • 适合需要多任务统一生成的科研与工业场景。

扩散模型在文本到图像生成中取得显著成功,激发了其向图像理解、操作和感知等多模态任务扩展的兴趣。这些任务需要在视觉与文本模态间具备高级语义理解能力,尤其是在处理复杂语义指令时。然而,现有方法通常依赖视觉语言模型(VLMs)或模块化设计进行语义引导,导致架构碎片化且计算效率低。为此,我们提出UniAlignment,一种基于单一扩散Transformer的统一多模态生成框架。该框架引入双流扩散训练策略,结合内在模态语义对齐与跨模态语义对齐,增强模型的跨模态一致性与指令遵循鲁棒性。此外,我们构建了SemGen-Bench,一个专为评估复杂文本指令下多模态语义一致性而设计的新基准。在多个任务与基准上的大量实验表明,UniAlignment显著优于现有基线,凸显了扩散模型在统一多模态生成中的巨大潜力。

原文摘要 · Abstract (English)

The remarkable success of diffusion models in text-to-image generation has sparked growing interest in expanding their capabilities to a variety of multi-modal tasks, including image understanding, manipulation, and perception. These tasks require advanced semantic comprehension across both visual and textual modalities, especially in scenarios involving complex semantic instructions. However, existing approaches often rely heavily on vision-language models (VLMs) or modular designs for semantic guidance, leading to fragmented architectures and computational inefficiency. To address these challenges, we propose UniAlignment, a unified multimodal generation framework within a single diffusion transformer. UniAlignment introduces a dual-stream diffusion training strategy that incorporates both intrinsic-modal semantic alignment and cross-modal semantic alignment, thereby enhancing the model's cross-modal consistency and instruction-following robustness. Additionally, we present SemGen-Bench, a new benchmark specifically designed to evaluate multimodal semantic consistency under complex textual instructions. Extensive experiments across multiple tasks and benchmarks demonstrate that UniAlignment outperforms existing baselines, underscoring the significant potential of diffusion models in unified multimodal generation.

多模态生成扩散模型语义对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。