用思维链推理提升复杂道路场景的异常检测能力
CoT-Segmenter: Enhancing OOD Detection in Dense Road Scenes via Chain-of-Thought Reasoning
- 通过思维链提示让大模型分析图像并推理异常
- 在密集、远距离、大物体场景下显著优于现有方法
- 适合自动驾驶安全系统研发人员参考
有效的分布外(OOD)检测对保障语义分割模型在复杂道路环境中的可靠性至关重要,尤其在安全性与准确性要求极高的驾驶场景中。尽管大型语言模型(如GPT-4)通过思维链(CoT)提示显著提升了多模态推理能力,但基于CoT的视觉推理在OOD语义分割中的应用仍处于空白。本文通过对道路异常场景的深入分析,识别出三类当前最先进的方法普遍难以应对的挑战:(1)物体密集重叠;(2)远距离小物体;(3)前景主导的大目标。为此,我们提出一种新型基于思维链的框架,用于道路异常场景下的OOD检测。该方法利用GPT-4等基础模型的丰富知识与推理能力,通过与观察到的异常场景特征对齐的提示机制,提升图像理解与推理能力。大量实验表明,本框架在标准基准及我们新定义的RoadAnomaly数据集挑战子集上均持续优于现有最先进方法,为复杂驾驶环境中的OOD语义分割提供了鲁棒且可解释的解决方案。
原文摘要 · Abstract (English)
Effective Out-of-Distribution (OOD) detection is criti-cal for ensuring the reliability of semantic segmentation models, particularly in complex road environments where safety and accuracy are paramount. Despite recent advancements in large language models (LLMs), notably GPT-4, which significantly enhanced multimodal reasoning through Chain-of-Thought (CoT) prompting, the application of CoT-based visual reasoning for OOD semantic segmentation remains largely unexplored. In this paper, through extensive analyses of the road scene anomalies, we identify three challenging scenarios where current state-of-the-art OOD segmentation methods consistently struggle: (1) densely packed and overlapping objects, (2) distant scenes with small objects, and (3) large foreground-dominant objects. To address the presented challenges, we propose a novel CoT-based framework targeting OOD detection in road anomaly scenes. Our method leverages the extensive knowledge and reasoning capabilities of foundation models, such as GPT-4, to enhance OOD detection through improved image understanding and prompt-based reasoning aligned with observed problematic scene attributes. Extensive experiments show that our framework consistently outperforms state-of-the-art methods on both standard benchmarks and our newly defined challenging subset of the RoadAnomaly dataset, offering a robust and interpretable solution for OOD semantic segmentation in complex driving environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。