arXiv:2510.20819cs.CVcs.AI2025-10NeurIPS被引 5

无需对齐维度,用扩散模型实现任意模态间高效转换。

Towards General Modality Translation with Contrastive and Predictive Latent Diffusion Bridge

  • 在共享隐空间中构建跨模态桥梁,不依赖固定维度或特定架构。
  • 支持多视图转3D、图像超分等任务,在多个数据集上超越现有方法。
  • 适合研究通用跨模态生成的学者,尤其关注无监督翻译场景。

生成建模的最新进展使扩散模型成为从复杂数据分布采样的最先进工具。尽管这些模型在图像、音频等单模态领域表现卓越,但将其能力扩展至跨模态翻译(MT)仍面临挑战。现有方法常依赖于受限假设,如共享维度、高斯源先验和模态专用架构,限制了其通用性与理论基础。本文提出隐变量扩展的去噪扩散桥模型(LDDBM),一种通用跨模态翻译框架。通过在共享隐空间中学习任意模态间的映射,该方法无需对齐维度。引入对比对齐损失以保持配对样本的语义一致性,并设计无领域偏倚的编码器-解码器结构,专用于隐空间噪声预测。此外,提出预测损失引导训练向准确跨域转换,并探索多种训练策略提升稳定性。该方法支持任意模态对,且在多视图转3D、图像超分辨率、多视图场景合成等多样任务中表现强劲。大量实验与消融分析验证了框架有效性,建立了一项新的强基线。

原文摘要 · Abstract (English)

Recent advances in generative modeling have positioned diffusion models as state-of-the-art tools for sampling from complex data distributions. While these models have shown remarkable success across single-modality domains such as images and audio, extending their capabilities to Modality Translation (MT), translating information across different sensory modalities, remains an open challenge. Existing approaches often rely on restrictive assumptions, including shared dimensionality, Gaussian source priors, and modality-specific architectures, which limit their generality and theoretical grounding. In this work, we propose the Latent Denoising Diffusion Bridge Model (LDDBM), a general-purpose framework for modality translation based on a latent-variable extension of Denoising Diffusion Bridge Models. By operating in a shared latent space, our method learns a bridge between arbitrary modalities without requiring aligned dimensions. We introduce a contrastive alignment loss to enforce semantic consistency between paired samples and design a domain-agnostic encoder-decoder architecture tailored for noise prediction in latent space. Additionally, we propose a predictive loss to guide training toward accurate cross-domain translation and explore several training strategies to improve stability. Our approach supports arbitrary modality pairs and performs strongly on diverse MT tasks, including multi-view to 3D shape generation, image super-resolution, and multi-view scene synthesis. Comprehensive experiments and ablations validate the effectiveness of our framework, establishing a new strong baseline in general modality translation. For more information, see our project page: https://sites.google.com/view/lddbm/home.

跨模态翻译扩散模型隐空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。