提出新模型CoVAE,更好保留多模态数据相关性。
CoVAE: correlated multimodal generative modeling
- 设计新架构,显式建模多模态间相关性
- 在真实与合成数据上实现精准跨模态重建
- 适合需要准确不确定性估计的研究者
多模态变分自编码器已成为从丰富多模态数据中提取有效表示的常用工具。然而,这类模型依赖于潜在空间中的融合策略,破坏了多模态数据的联合统计结构,对生成效果和不确定性量化产生深远影响。本文提出相关变分自编码器(CoVAE),一种能捕捉模态间相关性的新型生成架构。我们在多个真实和合成数据集上测试CoVAE,结果表明其不仅能实现准确的跨模态重构,还能有效量化相关不确定性。
原文摘要 · Abstract (English)
Multimodal Variational Autoencoders have emerged as a popular tool to extract effective representations from rich multimodal data. However, such models rely on fusion strategies in latent space that destroy the joint statistical structure of the multimodal data, with profound implications for generation and uncertainty quantification. In this work, we introduce Correlated Variational Autoencoders (CoVAE), a new generative architecture that captures the correlations between modalities. We test CoVAE on a number of real and synthetic data sets demonstrating both accurate cross-modal reconstruction and effective quantification of the associated uncertainties.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。