arXiv:2504.05908cs.CVcs.AI2025-04CVPR被引 6

用概率推理提升自动驾驶对不确定场景的理解能力

PRIMEDrive-CoT: A Precognitive Chain-of-Thought Framework for Uncertainty-Aware Object Interaction in Driving Scene Scenario

  • 融合激光雷达与多视角图像,通过贝叶斯图网络建模不确定性
  • 在DriveCoT数据集上优于现有最先进模型,决策更可靠
  • 适合关注自动驾驶安全与可解释性的研究者

驾驶场景理解是自动驾驶中的关键现实问题,涉及对车辆、行人和交通信号等环境元素的解析与关联。尽管自动驾驶技术不断进步,传统方法依赖确定性模型,难以捕捉真实驾驶中固有的不确定性。为此,我们提出PRIMEDrive-CoT,一种面向驾驶场景下对象交互与思维链(CoT)推理的新型不确定性感知模型。该方法结合基于激光雷达的3D目标检测与多视角RGB图像,确保场景理解的可解释性与可靠性。利用贝叶斯图神经网络(BGNNs)建模不确定性与风险评估,并结合对象动态与上下文线索进行概率推理,同时通过思维链生成可解释决策,Grad-CAM可视化注意力区域。在DriveCoT数据集上的大量实验表明,PRIMEDrive-CoT在性能上超越当前最先进的CoT与风险感知模型。

原文摘要 · Abstract (English)

Driving scene understanding is a critical real-world problem that involves interpreting and associating various elements of a driving environment, such as vehicles, pedestrians, and traffic signals. Despite advancements in autonomous driving, traditional pipelines rely on deterministic models that fail to capture the probabilistic nature and inherent uncertainty of real-world driving. To address this, we propose PRIMEDrive-CoT, a novel uncertainty-aware model for object interaction and Chain-of-Thought (CoT) reasoning in driving scenarios. In particular, our approach combines LiDAR-based 3D object detection with multi-view RGB references to ensure interpretable and reliable scene understanding. Uncertainty and risk assessment, along with object interactions, are modelled using Bayesian Graph Neural Networks (BGNNs) for probabilistic reasoning under ambiguous conditions. Interpretable decisions are facilitated through CoT reasoning, leveraging object dynamics and contextual cues, while Grad-CAM visualizations highlight attention regions. Extensive evaluations on the DriveCoT dataset demonstrate that PRIMEDrive-CoT outperforms state-of-the-art CoT and risk-aware models.

自动驾驶不确定性思维链贝叶斯

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。