用可控生成模型诊断航拍目标检测器,精准定位弱点并指导高效补数据。
Diagnosing Aerial-View Object Detectors with Foundational Image Generative Models

- 通过文本引导生成与属性编辑,构建可控制的合成测试环境。
- 合成场景下性能趋势与真实数据一致,弱项识别准确率高。
- 针对性补充少量真实数据,最多提升13%检测精度,效率远超盲目增广。
大规模图像生成模型的发展使得可控属性的逼真场景合成成为可能。除了数据增强外,其在航空与遥感视觉系统诊断中的潜力尚未被探索。本文提出一种面向航拍车辆检测的合成诊断框架,结合文本引导生成、属性可控编辑和自动化属性验证,构建可控制的合成测试环境。该框架能精细评估预训练检测器在多种场景类型和环境条件下的表现,这些条件在真实数据集中难以分离。在三个检测架构和三个真实航拍数据集上,合成场景的性能趋势与真实世界的弱点高度吻合。基于诊断结果,针对薄弱类别补充少量真实数据,检测平均精度(AP50)最高提升13%,且所需样本远少于非目标增强。结果表明,可控合成探针可预测真实域性能差距,并指导高效数据采集。该诊断框架模块化设计,未来可集成其他生成或视觉-语言模型。代码与数据集已公开:https://humansensinglab.github.io/AVODDiag/
原文摘要 · Abstract (English)
Recent advances in large-scale image generative models enable photorealistic scene synthesis with controllable attributes. Beyond data augmentation, their potential as diagnostic tools for trained vision systems remains unexplored in the aerial and remote sensing domains. We introduce a synthetic diagnostic framework for aerial-view vehicle detection that combines text-guided generation, attribute-controlled editing, and automated attribute verification to construct a controllable synthetic testbed. This enables fine-grained evaluation of pretrained detectors under diverse scene types and environmental conditions that are difficult to isolate in real datasets. Across three detection architectures and three real aerial datasets, synthetic scene-wise performance trends closely match real-world weaknesses. Guided by these diagnostics, targeted supplementation with small real datasets from the identified weak categories yields improvements of up to 13% AP50 while requiring substantially fewer additional samples than non-targeted augmentation. Our results show that controlled synthetic probing can predict real-domain performance gaps and guide efficient data collection. The proposed diagnostic framework is modular and can incorporate alternative generative or vision-language models as capabilities evolve. Our code and datasets are available here: https://humansensinglab.github.io/AVODDiag/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。