用扩散桥模型提升密集预测任务的精度与效率
DPBridge: Latent Diffusion Bridge for Dense Prediction
- 基于扩散桥框架,直接从输入图生成目标图,跳过噪声中间态
- 在深度估计等任务上超越现有方法,性能提升显著
- 适合需要高精度视觉回归的任务,如3D重建、图像修复
扩散模型在捕捉复杂数据分布方面表现出色,在诸多生成任务中取得优异成果。尽管其已被拓展至深度估计、表面法向预测等密集预测任务,但潜力尚未充分挖掘。由于目标信号图与输入图像像素对齐,传统噪声到数据的生成范式效率低下,而输入图像可作为比纯噪声更优的先验。扩散桥模型支持两类数据分布间的直接生成,是潜在替代方案,但通常无法利用大型预训练基础模型中的丰富视觉先验。为此,本文将扩散桥公式与结构化视觉先验结合,提出首个用于密集预测任务的潜空间扩散桥框架DPBridge。为解决扩散桥模型与预训练扩散主干之间的不兼容性,我们提出:(1) 可计算的反向转移核,支持最大似然训练;(2) 包含分布对齐归一化和图像一致性损失的微调策略。在多个基准测试上的实验验证了该方法在不同场景下均具优越性能,证明其有效性与泛化能力。
原文摘要 · Abstract (English)
Diffusion models demonstrate remarkable capabilities in capturing complex data distributions and have achieved compelling results in many generative tasks. While they have recently been extended to dense prediction tasks such as depth estimation and surface normal prediction, their full potential in this area remains underexplored. As target signal maps and input images are pixel-wise aligned, the conventional noise-to-data generation paradigm is inefficient, and input images can serve as a more informative prior compared to pure noise. Diffusion bridge models, which support data-to-data generation between two general data distributions, offer a promising alternative, but they typically fail to exploit the rich visual priors embedded in large pretrained foundation models. To address these limitations, we integrate diffusion bridge formulation with structured visual priors and introduce DPBridge, the first latent diffusion bridge framework for dense prediction tasks. To resolve the incompatibility between diffusion bridge models and pretrained diffusion backbones, we propose (1) a tractable reverse transition kernel for the diffusion bridge process, enabling maximum likelihood training scheme; (2) finetuning strategies including distribution-aligned normalization and image consistency loss. Experiments across extensive benchmarks validate that our method consistently achieves superior performance, demonstrating its effectiveness and generalization capability under different scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。