arXiv:2509.08628cs.CV2025-09被引 8

用部分配对数据实现跨域图像翻译,无需全量标注

LADB: Latent Aligned Diffusion Bridges for Semi-Supervised Domain Translation

  • 在共享隐空间中对齐源域与目标域分布
  • 深度到图像翻译任务上优于无配对方法,保持生成质量
  • 适合数据标注昂贵的现实场景,支持多源多目标翻译

扩散模型虽能生成高质量结果,但在数据稀缺域中面临挑战,需大量重训练或昂贵的成对数据。为此,我们提出隐空间对齐扩散桥(LADB),一种基于部分配对数据的样本到样本域翻译框架。通过在共享隐空间中对齐源域与目标域分布,LADB将预训练的源域扩散模型与在部分配对隐表示上训练的目标域隐空间扩散模型(LADM)无缝结合,实现确定性域映射,无需完全监督。相比无配对方法缺乏可控性、全配对方法依赖大规模领域特定数据,LADB通过混合配对与未配对隐空间耦合,在保真度与多样性间取得平衡。实验表明,其在部分监督下的深度到图像翻译任务中表现更优。进一步扩展至多源(深度图、分割掩码)和多目标类条件风格迁移任务,验证了其在多样异构场景中的泛化能力。最终,我们展示LADB作为可扩展、多功能的真实域翻译解决方案,尤其适用于标注成本高或不完整的场景。

原文摘要 · Abstract (English)

Diffusion models excel at generating high-quality outputs but face challenges in data-scarce domains, where exhaustive retraining or costly paired data are often required. To address these limitations, we propose Latent Aligned Diffusion Bridges (LADB), a semi-supervised framework for sample-to-sample translation that effectively bridges domain gaps using partially paired data. By aligning source and target distributions within a shared latent space, LADB seamlessly integrates pretrained source-domain diffusion models with a target-domain Latent Aligned Diffusion Model (LADM), trained on partially paired latent representations. This approach enables deterministic domain mapping without the need for full supervision. Compared to unpaired methods, which often lack controllability, and fully paired approaches that require large, domain-specific datasets, LADB strikes a balance between fidelity and diversity by leveraging a mixture of paired and unpaired latent-target couplings. Our experimental results demonstrate superior performance in depth-to-image translation under partial supervision. Furthermore, we extend LADB to handle multi-source translation (from depth maps and segmentation masks) and multi-target translation in a class-conditioned style transfer task, showcasing its versatility in handling diverse and heterogeneous use cases. Ultimately, we present LADB as a scalable and versatile solution for real-world domain translation, particularly in scenarios where data annotation is costly or incomplete.

扩散模型域翻译半监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。