提出Y型结构统一图文理解与生成,解决共享模型性能妥协问题。
UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation
- 采用浅层共享、深层分叉的Y型架构,分离理解与生成任务
- 在深层降低模态对齐以恢复空间细节,提升生成质量
- 性能媲美专用模型,适合多任务统一系统设计
统一图像理解和生成已成为多模态人工智能的前沿方向。尽管取得进展,统一模型的最佳架构仍不明确。本文分析了针对理解与生成的任务专用专家模型及现有统一模型的模态对齐行为,发现:理解任务随网络深度递增模态对齐,有助于语义信息积累;而生成任务则在浅层增强对齐,深层减弱对齐以恢复空间细节。这种差异导致全共享Transformer主干存在根本冲突,常引发任务间性能折损。为此,我们提出UniFork——一种Y型架构,浅层共享跨任务表示学习,深层采用任务专用分支,避免干扰。大量消融实验表明,UniFork持续优于传统全共享架构,性能达到或超越任务专用模型。
原文摘要 · Abstract (English)
Unified image understanding and generation has emerged as a promising paradigm in multimodal artificial intelligence. Despite recent progress, the optimal architectural design for such unified models remains an open challenge. In this work, we start by analyzing the modality alignment behaviors of task-specific expert models for understanding and generation, as well as current unified models. Our analysis reveals a crucial observation: understanding tasks benefit from a progressively increasing modality alignment across network depth, which helps build up semantic information for better comprehension; In contrast, generation tasks follow a different trend: modality alignment increases in the early layers but decreases in the deep layers to recover spatial details. These divergent alignment patterns create a fundamental conflict in fully shared Transformer backbones, where a uniform representational flow often leads to performance compromises across two tasks. Motivated by this finding, we introduce UniFork, a novel Y-shaped architecture that shares the shallow layers for cross-task representation learning, while employing task-specific branches in deeper layers to avoid task interference. This design effectively balances shared learning and task specialization. Through extensive ablation experiments, we demonstrate that Unifork consistently outperforms conventional fully shared Transformer architectures, and achieves performance on par with or better than task-specific models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。