arXiv:2412.05707cs.CV2024-12被引 4

用图像大模型的分段特征提升道路障碍物检测精度

Segment-Level Road Obstacle Detection Using Visual Foundation Model Priors and Likelihood Ratios

  • 以图像块为单位而非像素,利用大模型语义特征
  • 在两个数据集上达到当前最优效果,无需调阈值
  • 适合自动驾驶系统对准确率和鲁棒性的高要求

道路障碍物检测对自动驾驶安全至关重要。现有方法通常对每个像素打分并设阈值生成结果,但阈值选择困难,且像素级分类易产生碎片化预测和大量误检。本文提出一种新方法,利用视觉基础模型的分段特征与似然比,直接预测道路障碍物。通过聚焦于图像片段而非单个像素,该方法提升了检测精度,减少了误报,并增强了对场景变化的鲁棒性。我们在RoadObstacle和LostAndFound数据集上进行了基准测试,结果表明该方法在无需预设阈值的情况下达到了当前最优性能。

原文摘要 · Abstract (English)

Detecting road obstacles is essential for autonomous vehicles to navigate dynamic and complex traffic environments safely. Current road obstacle detection methods typically assign a score to each pixel and apply a threshold to generate final predictions. However, selecting an appropriate threshold is challenging, and the per-pixel classification approach often leads to fragmented predictions with numerous false positives. In this work, we propose a novel method that leverages segment-level features from visual foundation models and likelihood ratios to predict road obstacles directly. By focusing on segments rather than individual pixels, our approach enhances detection accuracy, reduces false positives, and offers increased robustness to scene variability. We benchmark our approach against existing methods on the RoadObstacle and LostAndFound datasets, achieving state-of-the-art performance without needing a predefined threshold.

障碍物检测视觉模型自动驾驶分段识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。