arXiv:2606.13381cs.LG2026-06中稿 · ICML

提升多模态生成中质量与一致性平衡,让图像文字更协调真实。

Hölder++: Improving the Quality-Coherence Trade-off in Multimodal VAEs

论文配图:Hölder++: Improving the Quality-Coherence Trade-off in Multimodal VAEs
图 1 · 摘自论文原文
  • 采用精确的霍尔德池化机制,不依赖近似方法。
  • 分离共享与私有表征,增强跨模态一致性与多样性。
  • 分层推理提升表征解耦,适合下游任务应用。

现有多模态变分自编码器(MMVAEs)在生成质量与跨模态一致性之间存在权衡——难以同时生成逼真、多样且语义一致的样本。近期工作表明,使用霍尔德池化的简单近似能提升一致性,但轻微牺牲多样性。受此启发,我们提出Hölder++:首次在多模态VAE中实现无近似的霍尔德池化;引入扩展架构,区分共享与私有(模态特定)表示(Hölder+);并设计分层推理,进一步增强共享与私有表示的解耦(Hölder++)。实验表明,Hölder++持续优化生成质量-一致性权衡,获得更结构化的潜在空间,并学习到对下游任务有信息量的共享表示。

原文摘要 · Abstract (English)

Existing approaches for multimodal variational autoencoders (VAEs) face a trade-off between generative quality and coherence-i.e., they struggle to generate realistic and diverse samples that, at the same time, are semantically consistent across modalities. A recent work shows that using a simple approximation to Hölder pooling as an aggregation method improves coherence over the SOTA MMVAE+, despite assuming a single shared representation across all modalities. Yet, it slightly compromises sample diversity. Inspired by this insight, we propose Hölder++, a novel multimodal VAE that improves the generative quality-coherence trade-off through: (i) the first implementation of Hölder pooling without any approximation for multimodal VAEs; (ii) an extended architecture that models distinct shared and private (i.e., modality-specific) representations (Hölder+); and (iii) hierarchical inference that further enhances the disentanglement between the shared and private representations (Hölder++). Our experiments corroborate that Hölder++ consistently improves the generative quality-coherence trade-off, yields more structured latent spaces, and learns shared representations that are informative for downstream tasks.

多模态生成变分自编码器表征学习霍尔德池化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。