arXiv:2411.02844cs.CVcs.AI2024-11被引 4

视觉显著性比深度估计更能提升目标检测性能。

Correlation of Object Detection Performance with Visual Saliency and Depth Estimation

  • 用显著性和深度预测模型分析检测精度关联性。
  • 显著性与检测精度相关性最高达0.459(Pascal VOC)。
  • 大物体相关性是小物体的3倍,适合针对性优化。

随着目标检测技术的发展,理解其与互补视觉任务的关系对优化模型架构和计算资源至关重要。本文研究了目标检测精度与两个基础视觉任务——深度预测和视觉显著性预测之间的相关性。通过在COCO和Pascal VOC数据集上使用DeepGaze IIE、Depth Anything、DPT-Large和Itti's模型进行综合实验,发现视觉显著性与目标检测精度的相关性始终强于深度预测(在Pascal VOC上最大mAρ达0.459,深度预测最高为0.283)。分析显示,不同物体类别间相关性差异显著,大物体的相关性可高达小物体的三倍。这表明将视觉显著性特征融入检测架构可能比深度信息更有效,尤其针对特定类别。此类类别特异性差异也为靶向特征工程和数据集设计提供了启示,有望实现更高效准确的目标检测系统。

原文摘要 · Abstract (English)

As object detection techniques continue to evolve, understanding their relationships with complementary visual tasks becomes crucial for optimising model architectures and computational resources. This paper investigates the correlations between object detection accuracy and two fundamental visual tasks: depth prediction and visual saliency prediction. Through comprehensive experiments using state-of-the-art models (DeepGaze IIE, Depth Anything, DPT-Large, and Itti's model) on COCO and Pascal VOC datasets, we find that visual saliency shows consistently stronger correlations with object detection accuracy (mA$ρ$ up to 0.459 on Pascal VOC) compared to depth prediction (mA$ρ$ up to 0.283). Our analysis reveals significant variations in these correlations across object categories, with larger objects showing correlation values up to three times higher than smaller objects. These findings suggest incorporating visual saliency features into object detection architectures could be more beneficial than depth information, particularly for specific object categories. The observed category-specific variations also provide insights for targeted feature engineering and dataset design improvements, potentially leading to more efficient and accurate object detection systems.

目标检测显著性深度估计性能优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。