用俯视图势场识别驾驶风险物,更准更快。
Potential Field as Scene Affordance for Behavior Change-Based Visual Risk Object Identification
- 用俯视图中的势场模拟场景可利用性,融合斥力与引力。
- 在RiskBench上空间/时间一致性提升20.3%和11.6%,nuScenes上分别提升5.4%和7.2%。
- 计算效率提高88%,适合自动驾驶系统实时风险检测。
我们研究行为变化驱动的视觉风险物体识别(Visual-ROI),这是智能驾驶系统中检测潜在危险的关键框架。现有方法在空间精度和时间一致性上存在显著局限,源于对场景可利用性的理解不完整。例如,常将不影响自身车辆的车辆误判为风险物体。此外,现有基于行为变化的方法效率低下,因因果推断在透视图像空间中执行。本文提出新框架,采用鸟瞰图(BEV)表示,利用势场作为场景可利用性建模,包括来自道路基础设施和交通参与者的排斥力,以及来自目标目的地的吸引力。通过在BEV语义分割结果基础上,根据语义标签分配不同能量水平来计算势场。我们在合成与真实世界数据集上进行了全面实验与消融研究,结果表明,在RiskBench数据集上空间一致性提升20.3%,时间一致性提升11.6%;在nuScenes数据集上空间精度提升5.4%,时间一致性提升7.2%。同时计算效率提升88%。
原文摘要 · Abstract (English)
We study behavior change-based visual risk object identification (Visual-ROI), a critical framework designed to detect potential hazards for intelligent driving systems. Existing methods often show significant limitations in spatial accuracy and temporal consistency, stemming from an incomplete understanding of scene affordance. For example, these methods frequently misidentify vehicles that do not impact the ego vehicle as risk objects. Furthermore, existing behavior change-based methods are inefficient because they implement causal inference in the perspective image space. We propose a new framework with a Bird's Eye View (BEV) representation to overcome the above challenges. Specifically, we utilize potential fields as scene affordance, involving repulsive forces derived from road infrastructure and traffic participants, along with attractive forces sourced from target destinations. In this work, we compute potential fields by assigning different energy levels according to the semantic labels obtained from BEV semantic segmentation. We conduct thorough experiments and ablation studies, comparing the proposed method with various state-of-the-art algorithms on both synthetic and real-world datasets. Our results show a notable increase in spatial and temporal consistency, with enhancements of 20.3% and 11.6% on the RiskBench dataset, respectively. Additionally, we can improve computational efficiency by 88%. We achieve improvements of 5.4% in spatial accuracy and 7.2% in temporal consistency on the nuScenes dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。