arXiv:2604.02097cs.CVcs.LG2026-04被引 4

LatentUM用统一潜在空间实现跨模态推理与生成,无需像素转换。

LatentUM: Unleashing the Potential of Interleaved Cross-Modal Reasoning via a Latent-Space Unified Model

  • 所有模态共享同一潜在语义空间,直接打通理解与生成路径。
  • 在视觉空间规划任务中达最新最好性能,自反思生成更逼真。
  • 适合需要多轮跨模态交互的AI系统开发,如智能体决策与世界建模。

统一模型(UMs)具备跨异构模态理解与生成的潜力。相比单纯生成视觉内容,通过交织式跨模态推理解决密集视觉思维问题、利用自我反思提升视觉生成质量,或基于逐步动作干预建模物理世界动态更具价值。然而,现有UMs因理解与生成采用分离的视觉表示,必须依赖像素解码作为桥梁,效率低下。本文提出LatentUM,将所有模态统一表示在共享语义潜在空间中,消除视觉理解与生成间的像素空间中介。该设计天然支持灵活的交织式跨模态推理与生成。除计算效率提升外,共享表示显著缓解编码器偏差,增强跨模态对齐,使LatentUM在视觉空间规划基准上达到当前最优表现,通过自反思推动视觉生成极限,并能在共享语义潜在空间中预测未来视觉状态,实现世界建模。

原文摘要 · Abstract (English)

Unified models (UMs) hold promise for their ability to understand and generate content across heterogeneous modalities. Compared to merely generating visual content, the use of UMs for interleaved cross-modal reasoning is more promising and valuable, e.g., for solving understanding problems that require dense visual thinking, improving visual generation through self-reflection, or modeling visual dynamics of the physical world guided by stepwise action interventions. However, existing UMs necessitate pixel decoding as a bridge due to their disjoint visual representations for understanding and generation, which is both ineffective and inefficient. In this paper, we introduce LatentUM, a novel unified model that represents all modalities within a shared semantic latent space, eliminating the need for pixel-space mediation between visual understanding and generation. This design naturally enables flexible interleaved cross-modal reasoning and generation. Beyond improved computational efficiency, the shared representation substantially alleviates codec bias and strengthens cross-modal alignment, allowing LatentUM to achieve state-of-the-art performance on the Visual Spatial Planning benchmark, push the limits of visual generation through self-reflection, and support world modeling by predicting future visual states within the shared semantic latent space.

统一模型跨模态推理潜在空间视觉生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。