提升自动驾驶在复杂城市路况下的感知与规划可靠性
A Reliable Context-Aware and Temporal Planning Framework for Autonomous Driving
- 通过质量评估与记忆机制筛选可靠感知特征
- 在nuScenes上实现61.5检测得分与52.3交并比
- 适合关注实时性与安全性的自动驾驶系统研发
自动驾驶车辆在密集城市交通中安全运行依赖于感知与规划的可靠性,尤其当车载传感器数据受损时。实际驾驶中,摄像头常受遮挡、运动模糊、光照变化和噪声干扰,若未加甄别地融合多帧数据,将导致轨迹规划不稳定,增加本车及周围道路使用者的碰撞风险。现有鸟瞰图(BEV)方法虽统一感知与规划,但大多无差别融合时间信息。本文提出可靠上下文感知与时间规划框架(RCT-AD),显式建模特征质量与时间一致性,以支持更安全、稳定的规划。其可靠上下文感知模块通过质量门控的先进先出(FILO)记忆机制,筛选可信特征并利用历史可靠信息重建受损观测,防止劣质输入破坏场景表示。时间轨迹规划器捕捉长期依赖与多智能体交互,生成更平滑、安全的轨迹;联合检测与分割头将语义与运动线索注入共享BEV空间,增强场景理解。在nuScenes基准测试中,RCT-AD超越近期端到端基线,在感知精度、运动预测与规划鲁棒性上均有提升,达到61.5的nuScenes检测得分、52.9的平均精确率与52.3的平均交并比,同时保持适用于实时部署的计算效率。
原文摘要 · Abstract (English)
Safe operation of autonomous vehicles in dense urban traffic depends on perception and planning that remain reliable when onboard sensing is degraded. In real driving conditions, camera observations are frequently corrupted by occlusion, motion blur, illumination change, and sensor noise, and when such degraded observations are aggregated indiscriminately over time, trajectory planning becomes unstable and collision risk rises for both the ego vehicle and surrounding road users. Recent Bird's-Eye-View (BEV) approaches unify perception and planning through a shared spatial representation, but most fuse temporal information across frames without assessing the reliability of the underlying observations. We present a Reliable Context-Aware and Temporal Planning framework for Autonomous Driving (RCT-AD) that explicitly models feature quality and temporal consistency to support safer, more consistent planning. A Reliable Context Awareness module scores per-frame reliability and selectively retains trustworthy features through a quality-gated First-In-Last-Out (FILO) memory mechanism, reconstructing degraded observations from reliable historical context so that corrupted inputs do not destabilize the scene representation. A Temporal Trajectory Planner captures long-term dependencies and multi-agent interactions to produce smoother, safety-aware trajectories, while a joint detection-and-segmentation head injects semantic and motion cues into the shared BEV space to strengthen scene understanding. Experiments on the nuScenes autonomous driving benchmark show that RCT-AD improves perception accuracy, motion prediction, and planning robustness over recent end-to-end baselines, achieving 61.5 nuScenes Detection Score, 52.9 mean Average Precision, and 52.3 mean Intersection over Union, while maintaining competitive computational efficiency suitable for real-time deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。