提出新型多模态自编码器,提升生成质量和表达能力。
Hellinger Multimodal Variational Autoencoders
- 基于赫林格距离构建概率融合机制,避免采样偏差
- 新增模态时隐空间表示更丰富,生成质量更优
- 适合需要多模态生成且追求高质量输出的研究者
多模态变分自编码器广泛用于弱监督生成学习。现有方法通常通过专家产品(PoE)、专家混合(MoE)或其组合来近似联合后验分布。本文从概率意见聚合视角重新审视多模态推断,基于 Hölder 池化中 α=0.5 的对称情形(对应 α-散度族中的唯一对称成员),导出一种矩匹配近似,称为赫林格近似。在此基础上,提出 HELVAE 模型,无需子采样,实现高效且有效的多模态建模:(i) 随着观察到更多模态,学习到更具表达力的潜在表示;(ii) 在生成连贯性与质量之间取得更好权衡,优于当前最优多模态 VAE 模型。
原文摘要 · Abstract (English)
Multimodal variational autoencoders (VAEs) are widely used for weakly supervised generative learning with multiple modalities. Predominant methods aggregate unimodal inference distributions using either a product of experts (PoE), a mixture of experts (MoE), or their combinations to approximate the joint posterior. In this work, we revisit multimodal inference through the lens of probabilistic opinion pooling, an optimization-based approach. We start from Hölder pooling with $α=0.5$, which corresponds to the unique symmetric member of the $α\text{-divergence}$ family, and derive a moment-matching approximation, termed Hellinger. We then leverage such an approximation to propose HELVAE, a multimodal VAE that avoids sub-sampling, yielding an efficient yet effective model that: (i) learns more expressive latent representations as additional modalities are observed; and (ii) empirically achieves better trade-offs between generative coherence and quality, outperforming state-of-the-art multimodal VAE models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。