用扩散模型生成风格化数据,提升目标检测在城市场景的泛化能力
Object Style Diffusion for Generalized Object Detection in Urban Scene
- 通过潜空间扩散模型生成带风格变化的伪目标域数据
- 在自动驾驶场景下实现超越现有方法的检测性能
- 可作为通用插件增强其他单域泛化方法
目标检测是计算机视觉的关键任务,广泛应用于自动驾驶和城市场景监控。然而,基于深度学习的方法通常需要大量标注数据,而这些数据在复杂多变的真实环境中获取成本高、难度大,严重制约了现有检测技术的泛化能力。为此,我们提出一种新型单域检测泛化方法GoDiff,利用预训练模型提升在未见域中的泛化性能。核心是伪目标域数据生成(PTDG)模块,该模块采用潜扩散模型生成保留源域特征但具有风格差异的伪数据,与源域数据结合以扩充训练集。同时引入跨风格实例归一化技术,融合PTDG生成的不同域风格特征,增强检测器鲁棒性。实验表明,该方法不仅显著提升现有检测器的泛化能力,还可作为即插即用的增强模块,适用于其他单域泛化方法,在自动驾驶场景中达到当前最优表现。
原文摘要 · Abstract (English)
Object detection is a critical task in computer vision, with applications in various domains such as autonomous driving and urban scene monitoring. However, deep learning-based approaches often demand large volumes of annotated data, which are costly and difficult to acquire, particularly in complex and unpredictable real-world environments. This dependency significantly hampers the generalization capability of existing object detection techniques. To address this issue, we introduce a novel single-domain object detection generalization method, named GoDiff, which leverages a pre-trained model to enhance generalization in unseen domains. Central to our approach is the Pseudo Target Data Generation (PTDG) module, which employs a latent diffusion model to generate pseudo-target domain data that preserves source domain characteristics while introducing stylistic variations. By integrating this pseudo data with source domain data, we diversify the training dataset. Furthermore, we introduce a cross-style instance normalization technique to blend style features from different domains generated by the PTDG module, thereby increasing the detector's robustness. Experimental results demonstrate that our method not only enhances the generalization ability of existing detectors but also functions as a plug-and-play enhancement for other single-domain generalization methods, achieving state-of-the-art performance in autonomous driving scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。