用压缩连续语义表示统一多模态理解与生成,提升图像编辑可控性。
UniCom: Unified Multimodal Modeling via Compressed Continuous Semantic Representations
- 通过注意力压缩器将密集特征提炼为紧凑统一表示
- 在生成任务中达到统一模型最优表现,收敛更快更稳定
- 无需VAE也能保持图像一致性,适合可控图像编辑
当前统一多模态模型通常依赖离散视觉分词器弥合模态差距,但离散化会损失细粒度语义信息,导致视觉理解性能不佳。相反,直接建模连续语义表示(如CLIP、SigLIP)在高维生成建模中面临收敛慢、训练不稳定的问题。为此,我们提出UniCom框架,通过压缩的连续表示统一多模态理解与生成。实验证明,降低通道维度比空间下采样更有效于重建与生成。我们设计了基于注意力的语义压缩器,将密集特征浓缩为紧凑统一表示。此外,验证了transfusion架构在收敛性与一致性上优于查询式设计。实验表明,UniCom在统一模型中达到最先进的生成性能。尤其通过保留丰富语义先验,实现了出色的图像编辑可控性,并在无VAE情况下仍保持图像一致性。
原文摘要 · Abstract (English)
Current unified multimodal models typically rely on discrete visual tokenizers to bridge the modality gap. However, discretization inevitably discards fine-grained semantic information, leading to suboptimal performance in visual understanding tasks. Conversely, directly modeling continuous semantic representations (e.g., CLIP, SigLIP) poses significant challenges in high-dimensional generative modeling, resulting in slow convergence and training instability. To resolve this dilemma, we introduce UniCom, a unified framework that harmonizes multimodal understanding and generation via compressed continuous representation. We empirically demonstrate that reducing channel dimension is significantly more effective than spatial downsampling for both reconstruction and generation. Accordingly, we design an attention-based semantic compressor to distill dense features into a compact unified representation. Furthermore, we validate that the transfusion architecture surpasses query-based designs in convergence and consistency. Experiments demonstrate that UniCom achieves state-of-the-art generation performance among unified models. Notably, by preserving rich semantic priors, it delivers exceptional controllability in image editing and maintains image consistency even without relying on VAE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。