arXiv:2601.06572cs.LGcs.AI2026-01中稿 · AISTATS 2026被引 2

提出新型多模态自编码器,提升生成质量和表达能力。

Hellinger Multimodal Variational Autoencoders

  • 基于赫林格距离构建概率融合机制,避免采样偏差
  • 新增模态时隐空间表示更丰富,生成质量更优
  • 适合需要多模态生成且追求高质量输出的研究者

多模态变分自编码器广泛用于弱监督生成学习。现有方法通常通过专家产品(PoE)、专家混合(MoE)或其组合来近似联合后验分布。本文从概率意见聚合视角重新审视多模态推断,基于 Hölder 池化中 α=0.5 的对称情形(对应 α-散度族中的唯一对称成员),导出一种矩匹配近似,称为赫林格近似。在此基础上,提出 HELVAE 模型,无需子采样,实现高效且有效的多模态建模:(i) 随着观察到更多模态,学习到更具表达力的潜在表示;(ii) 在生成连贯性与质量之间取得更好权衡,优于当前最优多模态 VAE 模型。

原文摘要 · Abstract (English)

Multimodal variational autoencoders (VAEs) are widely used for weakly supervised generative learning with multiple modalities. Predominant methods aggregate unimodal inference distributions using either a product of experts (PoE), a mixture of experts (MoE), or their combinations to approximate the joint posterior. In this work, we revisit multimodal inference through the lens of probabilistic opinion pooling, an optimization-based approach. We start from Hölder pooling with $α=0.5$, which corresponds to the unique symmetric member of the $α\text{-divergence}$ family, and derive a moment-matching approximation, termed Hellinger. We then leverage such an approximation to propose HELVAE, a multimodal VAE that avoids sub-sampling, yielding an efficient yet effective model that: (i) learns more expressive latent representations as additional modalities are observed; and (ii) empirically achieves better trade-offs between generative coherence and quality, outperforming state-of-the-art multimodal VAE models.

多模态变分自编码器生成模型概率融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。