arXiv:2508.17376cs.LGcs.CV2025-08被引 1

构建多模态共享潜在空间,提升跨模态生成质量与扩展性

ShaLa: Multimodal Shared Latent Space Modelling

  • 采用新型推理架构与扩散先验,优化共享潜在表示学习
  • 在多个基准上生成结果更连贯,质量优于当前最优的多模态VAE
  • 可拓展至更多模态,有效应对复杂共享空间建模挑战

本文提出一种新型生成框架——ShaLa,用于学习多模态数据间的共享潜在表示。现有先进方法通常关注模态特有细节的组合,易掩盖跨模态的高层语义概念。尽管低维潜在变量的多模态变分自编码器(Multimodal VAEs)旨在捕捉共享表示,以支持联合生成与跨模态推理,但其常面临表达能力不足的联合变分后验设计问题,导致生成质量较低。为此,ShaLa通过引入新型架构化推理模型与第二阶段表达性强的扩散先验,不仅提升了共享潜在表示的推断效率,还显著改善了下游多模态生成质量。我们在多个基准上对ShaLa进行了充分验证,结果显示其生成结果在连贯性与质量上均优于当前最先进多模态VAE。此外,相较于以往方法在处理高复杂度共享空间时的局限,ShaLa具备更强的模态扩展能力,能有效处理更多模态。

原文摘要 · Abstract (English)

This paper presents a novel generative framework for learning shared latent representations across multimodal data. Many advanced multimodal methods focus on capturing all combinations of modality-specific details across inputs, which can inadvertently obscure the high-level semantic concepts that are shared across modalities. Notably, Multimodal VAEs with low-dimensional latent variables are designed to capture shared representations, enabling various tasks such as joint multimodal synthesis and cross-modal inference. However, multimodal VAEs often struggle to design expressive joint variational posteriors and suffer from low-quality synthesis. In this work, ShaLa addresses these challenges by integrating a novel architectural inference model and a second-stage expressive diffusion prior, which not only facilitates effective inference of shared latent representation but also significantly improves the quality of downstream multimodal synthesis. We validate ShaLa extensively across multiple benchmarks, demonstrating superior coherence and synthesis quality compared to state-of-the-art multimodal VAEs. Furthermore, ShaLa scales to many more modalities while prior multimodal VAEs have fallen short in capturing the increasing complexity of the shared latent space.

多模态生成潜在空间变分自编码器扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。