提出新模型分离多模态数据中的共性与私有信息,提升生成质量。
Disentanglement of Variations with Multimodal Generative Modeling
- 用互信息正则化显式分离共享与私有变量
- 在挑战性数据集上实现更清晰的信息解耦和更好生成效果
- 适合研究多模态生成与表示学习的学者参考
多模态数据广泛存在于多个领域,学习其稳健表示对提升生成质量和下游任务性能至关重要。为应对不同模态间的异质性和关联性,现有方法通过两个独立变量提取共享信息与私有(模态特有)信息。尽管尝试强制两变量解耦,但在似然模型不足的挑战性数据集上仍表现不佳。本文提出信息解耦多模态变分自编码器(IDMVAE),引入基于互信息的严格正则化:通过跨视图互信息最大化提取共享变量,并采用循环一致性风格损失结合生成增强去除冗余。进一步引入扩散模型提升隐变量先验容量。这些组件相互补充。相比现有方法,IDMVAE在挑战性数据集上实现了共享与私有信息的更清晰分离,生成质量与语义连贯性均显著提升。
原文摘要 · Abstract (English)
Multimodal data are prevalent across various domains, and learning robust representations of such data is paramount to enhancing generation quality and downstream task performance. To handle heterogeneity and interconnections among different modalities, recent multimodal generative models extract shared and private (modality-specific) information with two separate variables. Despite attempts to enforce disentanglement between these two variables, these methods struggle with challenging datasets where the likelihood model is insufficient. In this paper, we propose Information-disentangled Multimodal VAE (IDMVAE) to explicitly address this issue, with rigorous mutual information-based regularizations, including cross-view mutual information maximization for extracting shared variables, and a cycle-consistency style loss for redundancy removal using generative augmentations. We further introduce diffusion models to improve the capacity of latent priors. These newly proposed components are complementary to each other. Compared to existing approaches, IDMVAE shows a clean separation between shared and private information, demonstrating superior generation quality and semantic coherence on challenging datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。