arXiv:2608.28603cs.AI2026-08

构建因果一致的跨模态语义空间,解决多模态模型生成不稳、语义漂移问题。

C3-UniMM: Causal Cycle-Consistent Unified Multimodal Modeling via Super Alignment and Shared Decoding Space

论文配图:C3-UniMM: Causal Cycle-Consistent Unified Multimodal Modeling via Super Alignment and Shared Decoding Space
图 1 · 摘自论文原文
  • 引入结构化潜在因果图作为共享语义空间,统一理解与生成任务。
  • 在多个任务上超越现有基线,尤其在组合泛化任务中提升显著。
  • 适合需要稳定跨模态生成与推理的研究者和开发者。

统一多模态模型旨在实现任意模态间的任意理解与生成。然而,现有方法主要依赖隐式统计相关性建模,缺乏跨模态结构一致性约束,导致语义漂移、组合泛化能力差及干预下不稳定等问题。本文提出C3-UniMM,基于因果循环一致性和超对齐的统一多模态建模框架。具体地,引入结构化潜在因果图(SLCG)作为共享跨模态语义空间,并设计统一的多模态编码模块,使理解和生成在相同的因果语义结构中协同优化。此外,提出统一解码空间,在跨模态生成过程中强制保持结构一致性与语义可逆性。理论分析表明,该方法显著提升了跨模态映射的可逆性与机制不变性。在多个理解、生成及组合泛化任务上的大量实验结果表明,C3-UniMM明显优于现有统一多模态基线。

原文摘要 · Abstract (English)

Unified Multimodal Models aim to achieve any-to-any understanding and generation across arbitrary modalities. However, existing methods primarily rely on modeling implicit statistical correlations and lack cross-modal structural consistency constraints. This deficiency leads to profound issues, including semantic drift, poor compositional generalization, and instability under interventions. In this paper, we propose C3-UniMM, a unified multimodal modeling framework based on Causal Cycle Consistency and Super Alignment. Specifically, we introduce a Structured Latent Causal Graph (SLCG) as a shared cross-modal semantic space and design unified multimodal encoding blocks, enabling understanding and generation to be synergistically optimized within the identical causal semantic structure. Furthermore, we propose a Unified Decoding Space to enforce structural preservation and semantic invertibility during the cross-modal generation process. Theoretical analyses demonstrate that our approach significantly enhances both the invertibility and mechanism invariance of cross-modal mappings. Extensive experimental results across multiple understanding, generation, and compositional generalization tasks indicate that C3-UniMM substantially outperforms existing unified multimodal baselines.

多模态因果建模统一框架生成稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。