arXiv:2502.03952cs.LGstat.ML2025-02被引 3

提出新模型提升多模态生成质量,避免传统方法的推理差距问题。

Bridging the inference gap in Mutimodal Variational Autoencoders

  • 分阶段训练:先用变分推断建模联合分布,再用归一化流建模条件分布。
  • 在多个基准数据集上达到当前最优生成效果,显著提升样本一致性。
  • 通过共享信息提取增强跨模态生成的连贯性,适合医疗、自动驾驶等场景。

从医学诊断到自动驾驶,关键应用依赖于异构多模态数据的融合。多模态变分自编码器提供了从已观测模态生成未观测模态的灵活且可扩展方法。近期采用专家混合聚合的模型存在理论局限,限制了其在复杂数据集上的生成质量。本文提出一种新型可解释模型,无需引入专家混合聚合即可学习联合与条件分布。模型采用多阶段训练流程:首先使用变分推断建模联合分布,随后利用归一化流(Normalizing Flows)建模条件分布,以更精确逼近真实后验。尤为重要的是,我们提出提取并利用模态间的共享信息,以提升生成样本的条件一致性。该方法在多个基准数据集上实现当前最优性能。

原文摘要 · Abstract (English)

From medical diagnosis to autonomous vehicles, critical applications rely on the integration of multiple heterogeneous data modalities. Multimodal Variational Autoencoders offer versatile and scalable methods for generating unobserved modalities from observed ones. Recent models using mixturesof-experts aggregation suffer from theoretically grounded limitations that restrict their generation quality on complex datasets. In this article, we propose a novel interpretable model able to learn both joint and conditional distributions without introducing mixture aggregation. Our model follows a multistage training process: first modeling the joint distribution with variational inference and then modeling the conditional distributions with Normalizing Flows to better approximate true posteriors. Importantly, we also propose to extract and leverage the information shared between modalities to improve the conditional coherence of generated samples. Our method achieves state-of-the-art results on several benchmark datasets.

多模态生成变分自编码器归一化流

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。