首个面向闭环自动驾驶的设备级鲁棒性评测基准,测试真实部署扰动下的系统稳定性。
Bench2Drive-Robust: Benchmarking Closed-Loop Autonomous Driving under Deployment Perturbations

- 构建三类部署扰动:摄像头帧丢失、自车状态误差、计算延迟,模拟真实硬件问题。
- 实测显示这些扰动使闭环驾驶性能显著下降,传统图像级测试无法覆盖此风险。
- 为端到端自动驾驶系统提供可复现的鲁棒性评估标准,适合研究部署可靠性者参考。
鲁棒性是自动驾驶系统在真实世界部署的关键要求。现有自动驾驶鲁棒性评测主要关注图像级退化(如恶劣天气或摄像头故障)对感知模块和开环规划的影响,但对部署中的系统级缺陷(如推理延迟、自车状态估计误差)在闭环端到端自动驾驶中的影响研究不足。这些缺陷会通过反馈回路累积并导致控制失稳。本文提出 Bench2Drive-Robust,据我们所知是首个面向闭环端到端自动驾驶的真实部署扰动设备级鲁棒性评测基准。系统评估来自三大来源的部署相关扰动:摄像头流失效(帧丢失、局部观测)、自车状态估计误差(GPS噪声、速度/里程计误差)以及计算引发的控制延迟(模型推理延迟)。评估了代表性端到端驾驶方法,并分析其在不同扰动强度下的鲁棒性表现。结果表明,这些部署相关扰动会显著降低闭环驾驶性能,暴露出传统图像级退化评测未涵盖的鲁棒性挑战。通过建立闭环评估协议并验证这些部署导向扰动的严重影响,Bench2Drive-Robust明确了端到端自动驾驶的实际鲁棒性问题,推动面向部署的鲁棒驾驶系统研究。
原文摘要 · Abstract (English)
Robustness is a critical requirement for deploying autonomous driving systems in the real world. Existing robustness benchmarks for autonomous driving have made important progress in studying the effects of image-level corruptions, such as adverse weather or camera degradation, on perception modules and open-loop planning outputs. However, deployment can also involve system-level imperfections, such as inference latency and ego-state estimation errors, which remain less studied in closed-loop E2E-AD evaluation. These imperfections can accumulate through the feedback loop and destabilize control. In this work, we present Bench2Drive-Robust, to our knowledge the first device-centric robustness benchmark for closed-loop end-to-end autonomous driving under realistic deployment perturbations. We systematically evaluate deployment-oriented perturbations arising from three major sources: camera-stream failures (frame drop, partial observation), ego-state estimation errors (GPS noise, and speed or odometry errors), and compute-induced control delay (model inference delay). We evaluate representative end-to-end driving methods and analyze their robustness under different perturbation severities. Our results show that these deployment-related perturbations can substantially degrade closed-loop driving performance, revealing robustness challenges that are not fully captured by conventional image-level corruption evaluations. By establishing a closed-loop evaluation protocol and demonstrating the substantial impact of these deployment-oriented perturbations, Bench2Drive-Robust defines practical robustness problems for end-to-end autonomous driving and encourages further research on deployment-aware robust driving systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。