arXiv:2412.19944cs.CV2024-12被引 6

零样本识别自动驾驶罕见危险,三阶段流程提升检测精度

Zero-shot Hazard Identification in Autonomous Driving: A Case Study on the COOOL Benchmark

  • 用边界框与光流变化点检测司机反应,分析驾驶异常行为
  • 结合邻近策略与ViT模型,实现对罕见危险物体的精准识别
  • 采用定制提示词的MOLMO模型生成上下文相关的危险描述

本文提交至COOOL竞赛,该基准用于检测和分类自动驾驶中的未标注危险。提出三阶段方法:(i)基于核函数的边界框与光流动态变化点检测,分析司机反应;(ii)结合朴素邻近策略与预训练ViT模型进行危险物体识别;(iii)使用MOLMO视觉语言模型,通过定制提示生成针对稀有、低分辨率危险的准确描述。所提流程显著优于基线方法,相对误差降低33%,在32支参赛队伍中排名第二。

原文摘要 · Abstract (English)

This paper presents our submission to the COOOL competition, a novel benchmark for detecting and classifying out-of-label hazards in autonomous driving. Our approach integrates diverse methods across three core tasks: (i) driver reaction detection, (ii) hazard object identification, and (iii) hazard captioning. We propose kernel-based change point detection on bounding boxes and optical flow dynamics for driver reaction detection to analyze motion patterns. For hazard identification, we combined a naive proximity-based strategy with object classification using a pre-trained ViT model. At last, for hazard captioning, we used the MOLMO vision-language model with tailored prompts to generate precise and context-aware descriptions of rare and low-resolution hazards. The proposed pipeline outperformed the baseline methods by a large margin, reducing the relative error by 33%, and scored 2nd on the final leaderboard consisting of 32 teams.

自动驾驶零样本危险检测视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。