用生成模型合成航拍图提升车辆检测跨域能力
Adapting Vehicle Detectors for Aerial Imagery to Unseen Domains with Weak Supervision
- 用微调的扩散模型生成逼真航拍图像和标签进行数据增强
- 在多个新域上检测准确率提升7%至40%,最高超基准50%以上
- 适合做遥感、交通监控等跨域检测任务的研究者参考
航拍图像中的车辆检测在交通监控、城市规划和国防情报中至关重要。深度学习虽已取得最优效果,但训练数据所在地理区域与目标区域存在环境、道路布局、车种、成像参数(如分辨率、光照、角度)差异,导致性能显著下降。本文提出一种新方法,利用生成式AI合成高质量航拍图像及标注,通过数据增强改进检测器训练。核心贡献是构建多阶段、多模态知识迁移框架,采用微调的潜在扩散模型(LDMs)缩小源域与目标域分布差距。在多种航拍图像域上的实验表明,该方法在AP50上优于仅使用源域监督学习、弱监督适应、无监督域自适应及开集检测器4%-23%、6%-10%、7%-40%和超过50%。此外,本文还发布了来自新西兰和犹他州的新标注航拍数据集,以支持该领域研究。项目页面:https://humansensinglab.github.io/AGenDA
原文摘要 · Abstract (English)
Detecting vehicles in aerial imagery is a critical task with applications in traffic monitoring, urban planning, and defense intelligence. Deep learning methods have provided state-of-the-art (SOTA) results for this application. However, a significant challenge arises when models trained on data from one geographic region fail to generalize effectively to other areas. Variability in factors such as environmental conditions, urban layouts, road networks, vehicle types, and image acquisition parameters (e.g., resolution, lighting, and angle) leads to domain shifts that degrade model performance. This paper proposes a novel method that uses generative AI to synthesize high-quality aerial images and their labels, improving detector training through data augmentation. Our key contribution is the development of a multi-stage, multi-modal knowledge transfer framework utilizing fine-tuned latent diffusion models (LDMs) to mitigate the distribution gap between the source and target environments. Extensive experiments across diverse aerial imagery domains show consistent performance improvements in AP50 over supervised learning on source domain data, weakly supervised adaptation methods, unsupervised domain adaptation methods, and open-set object detectors by 4-23%, 6-10%, 7-40%, and more than 50%, respectively. Furthermore, we introduce two newly annotated aerial datasets from New Zealand and Utah to support further research in this field. Project page is available at: https://humansensinglab.github.io/AGenDA
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。