arXiv:2412.20487cs.LGcs.CV2024-12AAAI被引 10

用几何均值思想改进多模态生成模型,更准确捕捉跨模态信息。

Multimodal Variational Autoencoder: a Barycentric View

  • 以测地线均值重构多模态表示,统一了传统专家模型的框架
  • 在三个基准上优于经典方法,尤其在缺失模态时表现更稳定
  • 适合处理不完整多模态数据的研究者使用

现实世界现象常包含视觉、声音等多种信号模态。近年来,生成模型尤其是变分自编码器(VAE)在多模态表征学习中受到关注,尤其针对模态缺失情况。这类模型旨在学习既共性又个性的表征。以往工作多基于专家混合(MoE)或专家乘积(PoE)框架,通过加权KL散度聚合单模态后验分布。本文提出一种新的通用理论框架:将多模态建模视为测地线均值(barycenter)问题,证明PoE与MoE是特定加权不对称KL散度下的特例。进一步,引入基于2-Wasserstein距离的测地线均值,能更好保留单模态分布的几何结构,同时捕捉模态共享与特异性信息。在三个多模态基准上的实验证明该方法有效,性能优于现有方法。

原文摘要 · Abstract (English)

Multiple signal modalities, such as vision and sounds, are naturally present in real-world phenomena. Recently, there has been growing interest in learning generative models, in particular variational autoencoder (VAE), to for multimodal representation learning especially in the case of missing modalities. The primary goal of these models is to learn a modality-invariant and modality-specific representation that characterizes information across multiple modalities. Previous attempts at multimodal VAEs approach this mainly through the lens of experts, aggregating unimodal inference distributions with a product of experts (PoE), a mixture of experts (MoE), or a combination of both. In this paper, we provide an alternative generic and theoretical formulation of multimodal VAE through the lens of barycenter. We first show that PoE and MoE are specific instances of barycenters, derived by minimizing the asymmetric weighted KL divergence to unimodal inference distributions. Our novel formulation extends these two barycenters to a more flexible choice by considering different types of divergences. In particular, we explore the Wasserstein barycenter defined by the 2-Wasserstein distance, which better preserves the geometry of unimodal distributions by capturing both modality-specific and modality-invariant representations compared to KL divergence. Empirical studies on three multimodal benchmarks demonstrated the effectiveness of the proposed method.

多模态变分自编码器测地线均值生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。