arXiv:2505.01134cs.LG2025-05ICML被引 3

提出新方法联合多模态数据分布,提升生成质量与一致性。

Aggregation of Dependent Expert Distributions in Multimodal Variational Autoencoders

  • 用依赖专家共识机制融合单模态分布,避免独立假设。
  • 生成质量接近单模态模型,且随模态增加更稳定。
  • 适合多模态生成与分类任务,尤其关注生成一致性。

多模态变分自编码器(VAE)需估计联合分布以计算证据下界(ELBO)。现有方法如专家乘积和混合,为简化假设各模态独立,这过于乐观。本文提出一种基于依赖专家共识(CoDE)的新聚合方法,无需独立性假设。利用CoDE,我们设计新型ELBO,通过学习各模态子集的贡献来近似多模态数据的联合似然。所提CoDE-VAE模型在生成连贯性与质量之间取得更好平衡,且生成对数似然估计更精确。随着模态数量增加,其生成质量差距显著减小,在某些情况下达到与单模态VAE相当的水平,这是多数现有方法缺失的优势。此外,分类准确率与当前最优多模态VAE模型相当。

原文摘要 · Abstract (English)

Multimodal learning with variational autoencoders (VAEs) requires estimating joint distributions to evaluate the evidence lower bound (ELBO). Current methods, the product and mixture of experts, aggregate single-modality distributions assuming independence for simplicity, which is an overoptimistic assumption. This research introduces a novel methodology for aggregating single-modality distributions by exploiting the principle of consensus of dependent experts (CoDE), which circumvents the aforementioned assumption. Utilizing the CoDE method, we propose a novel ELBO that approximates the joint likelihood of the multimodal data by learning the contribution of each subset of modalities. The resulting CoDE-VAE model demonstrates better performance in terms of balancing the trade-off between generative coherence and generative quality, as well as generating more precise log-likelihood estimations. CoDE-VAE further minimizes the generative quality gap as the number of modalities increases. In certain cases, it reaches a generative quality similar to that of unimodal VAEs, which is a desirable property that is lacking in most current methods. Finally, the classification accuracy achieved by CoDE-VAE is comparable to that of state-of-the-art multimodal VAE models.

多模态变分自编码器生成模型联合分布

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。