提升多模态生成中质量与一致性平衡,让图像文字更协调真实。
Hölder++: Improving the Quality-Coherence Trade-off in Multimodal VAEs

- 采用精确的霍尔德池化机制,不依赖近似方法。
- 分离共享与私有表征,增强跨模态一致性与多样性。
- 分层推理提升表征解耦,适合下游任务应用。
现有多模态变分自编码器(MMVAEs)在生成质量与跨模态一致性之间存在权衡——难以同时生成逼真、多样且语义一致的样本。近期工作表明,使用霍尔德池化的简单近似能提升一致性,但轻微牺牲多样性。受此启发,我们提出Hölder++:首次在多模态VAE中实现无近似的霍尔德池化;引入扩展架构,区分共享与私有(模态特定)表示(Hölder+);并设计分层推理,进一步增强共享与私有表示的解耦(Hölder++)。实验表明,Hölder++持续优化生成质量-一致性权衡,获得更结构化的潜在空间,并学习到对下游任务有信息量的共享表示。
原文摘要 · Abstract (English)
Existing approaches for multimodal variational autoencoders (VAEs) face a trade-off between generative quality and coherence-i.e., they struggle to generate realistic and diverse samples that, at the same time, are semantically consistent across modalities. A recent work shows that using a simple approximation to Hölder pooling as an aggregation method improves coherence over the SOTA MMVAE+, despite assuming a single shared representation across all modalities. Yet, it slightly compromises sample diversity. Inspired by this insight, we propose Hölder++, a novel multimodal VAE that improves the generative quality-coherence trade-off through: (i) the first implementation of Hölder pooling without any approximation for multimodal VAEs; (ii) an extended architecture that models distinct shared and private (i.e., modality-specific) representations (Hölder+); and (iii) hierarchical inference that further enhances the disentanglement between the shared and private representations (Hölder++). Our experiments corroborate that Hölder++ consistently improves the generative quality-coherence trade-off, yields more structured latent spaces, and learns shared representations that are informative for downstream tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。