arXiv:2412.11807cs.CVcs.AI2024-12AAAI被引 24

用物理模型模拟真实光照干扰,提升目标检测在未知场景的泛化能力

PhysAug: A Physical-guided and Frequency-based Data Augmentation for Single-Domain Generalized Object Detection

  • 基于大气光学原理构建频域扰动模型,模拟真实成像条件
  • 在DWD和Cityscape-C上分别提升7.3%和7.2%的检测性能
  • 无需改动网络结构,适合追求鲁棒性的实际部署场景

单领域广义目标检测(S-DGOD)旨在仅使用单一源域数据训练,以实现对多种未见目标域的鲁棒检测。现有方法多依赖视觉变换组合进行数据增强,但缺乏真实世界先验知识,限制了训练数据分布的多样性。为此,本文提出PhysAug,一种基于物理模型的非理想成像条件数据增强方法。结合大气光学原理,构建通用扰动模型,利用图像频谱模拟光线与大气粒子相互作用带来的视觉退化。该方法使检测器学习域不变表征,显著提升跨域泛化能力。不改变网络结构或损失函数,已在多个S-DGOD数据集上超越现有最优方法,在DWD和Cityscape-C上分别取得7.3%和7.2%的性能提升,验证了其在真实场景下的有效性。

原文摘要 · Abstract (English)

Single-Domain Generalized Object Detection~(S-DGOD) aims to train on a single source domain for robust performance across a variety of unseen target domains by taking advantage of an object detector. Existing S-DGOD approaches often rely on data augmentation strategies, including a composition of visual transformations, to enhance the detector's generalization ability. However, the absence of real-world prior knowledge hinders data augmentation from contributing to the diversity of training data distributions. To address this issue, we propose PhysAug, a novel physical model-based non-ideal imaging condition data augmentation method, to enhance the adaptability of the S-DGOD tasks. Drawing upon the principles of atmospheric optics, we develop a universal perturbation model that serves as the foundation for our proposed PhysAug. Given that visual perturbations typically arise from the interaction of light with atmospheric particles, the image frequency spectrum is harnessed to simulate real-world variations during training. This approach fosters the detector to learn domain-invariant representations, thereby enhancing its ability to generalize across various settings. Without altering the network architecture or loss function, our approach significantly outperforms the state-of-the-art across various S-DGOD datasets. In particular, it achieves a substantial improvement of $7.3\%$ and $7.2\%$ over the baseline on DWD and Cityscape-C, highlighting its enhanced generalizability in real-world settings.

目标检测数据增强域泛化物理建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。