测试开放词汇目标检测在低质量图像下的表现,发现模型对严重退化敏感。
Evaluating the Performance of Open-Vocabulary Object Detection in Low-quality Image
- 构建真实世界低质量图像模拟数据集,评估模型鲁棒性。
- 高阶退化下所有模型性能骤降,但OWLv2表现最稳定。
- 适合关注视觉模型泛化能力的研究者参考。
开放词汇目标检测使模型能够定位和识别预定义类别之外的物体,有望达到接近人类的识别能力。本研究旨在评估现有模型在低质量图像条件下的开放词汇目标检测性能。为此,我们引入了一个模拟真实世界低质量图像的新数据集。实验结果表明,尽管在低级别图像退化下,开放词汇目标检测模型的mAP得分无显著下降,但在高级别图像退化下,所有模型性能均急剧下降。其中,OWLv2模型在各类退化条件下均表现出更优的稳定性,而OWL-ViT、GroundingDINO和Detic则出现显著性能下降。我们将公开数据集和代码,以促进后续研究。
原文摘要 · Abstract (English)
Open-vocabulary object detection enables models to localize and recognize objects beyond a predefined set of categories and is expected to achieve recognition capabilities comparable to human performance. In this study, we aim to evaluate the performance of existing models on open-vocabulary object detection tasks under low-quality image conditions. For this purpose, we introduce a new dataset that simulates low-quality images in the real world. In our evaluation experiment, we find that although open-vocabulary object detection models exhibited no significant decrease in mAP scores under low-level image degradation, the performance of all models dropped sharply under high-level image degradation. OWLv2 models consistently performed better across different types of degradation, while OWL-ViT, GroundingDINO, and Detic showed significant performance declines. We will release our dataset and codes to facilitate future studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。