用马尔可夫链优化多模态生成,让不同模态更协调。
Learning Multimodal Energy-Based Model with Multimodal Variational Auto-Encoder via MCMC Revision

- 通过交替优化生成器与推断模型,结合马尔可夫链采样提升多模态建模
- 在多个数据集上实现比基线更真实、更一致的多模态合成结果
- 适合需要高质量多模态生成的研究者,如跨模态图像-文本生成
能量模型(EBM)是灵活的深度生成模型,适合捕捉多模态数据中的复杂依赖关系。然而,基于最大似然学习多模态EBM需在联合数据空间中进行马尔可夫链蒙特卡洛(MCMC)采样,噪声初始化的朗之万动力学常混杂不良,难以发现一致的跨模态关联。多模态变分自编码器(VAE)通过共享潜在生成器和联合推理模型,在捕捉此类依赖方面取得进展,但二者均参数化为单模态高斯(或拉普拉斯)分布,严重限制其对多模态数据复杂结构的逼近能力。本文研究多模态EBM、共享潜在生成器与联合推理模型的学习问题,提出一个将它们的最大似然更新与数据空间和潜在空间中的对应MCMC修正相交织的学习框架。具体而言,生成器被训练以产生能作为EBM采样的良好初始状态的连贯多模态样本;而推理模型则被训练以提供生成器后验采样的信息性潜在初始状态。两者互为补充,共同促进有效EBM采样与学习,生成真实且连贯的多模态样本。大量实验表明,该方法在多模态合成质量与一致性上优于多种基线。我们还进行了多维度分析与消融实验,验证了所提框架的有效性与可扩展性。
原文摘要 · Abstract (English)
Energy-based models (EBMs) are a flexible class of deep generative models and are well-suited to capture complex dependencies in multimodal data. However, learning multimodal EBM by maximum likelihood requires Markov Chain Monte Carlo (MCMC) sampling in the joint data space, where noise-initialized Langevin dynamics often mixes poorly and fails to discover coherent inter-modal relationships. Multimodal VAEs have made progress in capturing such inter-modal dependencies by introducing a shared latent generator and a joint inference model. However, both the shared latent generator and joint inference model are parameterized as unimodal Gaussian (or Laplace), which severely limits their ability to approximate the complex structure induced by multimodal data. In this work, we study the learning problem of the multimodal EBM, shared latent generator, and joint inference model. We present a learning framework that effectively interweaves their MLE updates with corresponding MCMC refinements in both the data and latent spaces. Specifically, the generator is learned to produce coherent multimodal samples that serve as strong initial states for EBM sampling, while the inference model is learned to provide informative latent initializations for generator posterior sampling. Together, these two models serve as complementary models that enable effective EBM sampling and learning, yielding realistic and coherent multimodal EBM samples. Extensive experiments demonstrate superior performance for multimodal synthesis quality and coherence compared to various baselines. We conduct various analyses and ablation studies to validate the effectiveness and scalability of the proposed multimodal framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。