用扩散模型同时做图像生成和密集感知,提升效果与效率
Diff-2-in-1: Bridging Generation and Dense Perception with Diffusion Models
- 利用扩散去噪过程统一生成与感知任务
- 生成真实且有用的多模态数据,提升感知性能
- 自优化机制让生成数据持续改进感知效果
除了高保真图像合成,扩散模型近期在密集视觉感知任务中也展现出良好表现。然而,现有方法大多将扩散模型作为独立组件,仅用于现成的数据增强或特征提取,未能充分发挥其潜力。本文提出统一、通用的扩散框架 Diff-2-in-1,通过独特地利用扩散去噪过程,实现多模态数据生成与密集视觉感知的协同处理。该框架进一步通过去噪网络生成符合原始训练集分布的多模态数据,增强判别性感知能力。重要的是,Diff-2-in-1 采用新颖的自改进学习机制,优化生成数据的多样性和真实性,从而提升感知性能。大量实验验证了该框架的有效性,在多种判别性骨干网络上均取得一致性能提升,并能生成高质量、具实用性的多模态数据。
原文摘要 · Abstract (English)
Beyond high-fidelity image synthesis, diffusion models have recently exhibited promising results in dense visual perception tasks. However, most existing work treats diffusion models as a standalone component for perception tasks, employing them either solely for off-the-shelf data augmentation or as mere feature extractors. In contrast to these isolated and thus sub-optimal efforts, we introduce a unified, versatile, diffusion-based framework, Diff-2-in-1, that can simultaneously handle both multi-modal data generation and dense visual perception, through a unique exploitation of the diffusion-denoising process. Within this framework, we further enhance discriminative visual perception via multi-modal generation, by utilizing the denoising network to create multi-modal data that mirror the distribution of the original training set. Importantly, Diff-2-in-1 optimizes the utilization of the created diverse and faithful data by leveraging a novel self-improving learning mechanism. Comprehensive experimental evaluations validate the effectiveness of our framework, showcasing consistent performance improvements across various discriminative backbones and high-quality multi-modal data generation characterized by both realism and usefulness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。