首个面向自动驾驶的视觉语言模型鲁棒性评测基准,提出更可靠的端到端驾驶框架。
RoboDriveVLM: A Novel Benchmark and Baseline towards Robust Vision-Language Models for Autonomous Driving
- 构建11种模拟场景的鲁棒性测试集,覆盖传感器与提示双重干扰。
- 在64,559个轨迹预测案例中验证现有系统易受环境与数据干扰影响。
- 通过多模态融合与跨模态知识蒸馏提升模型在真实场景下的稳定性。
当前基于视觉语言模型(VLM)的端到端自动驾驶系统常直接利用大语言模型根据场景理解生成驾驶决策,但在真实场景中存在多重风险。为评估VLM在自动驾驶中的实际可行性,我们提出RoboDriveBench——首个专注于端到端轨迹预测任务的鲁棒性基准。该基准通过11种模拟场景系统评估两类现实挑战:6类由环境变化引发的传感器扰动,以及5类由人为干预和数据传输故障导致的提示扰动。每类扰动包含250个独特驾驶场景和5,689帧图像,总计64,559个轨迹预测案例。为应对这些挑战,我们提出新型自动驾驶框架RoboDriveVLM,通过将激光雷达、雷达等多模态数据映射至统一潜在空间以增强鲁棒性,并引入基于跨模态知识蒸馏的测试时自适应(TTA)方法。大量实验揭示了现有VLM系统在真实部署中的局限性,同时提供了更可靠解决方案。代码与数据集将公开发布。
原文摘要 · Abstract (English)
Current Vision-Language Model (VLM)-based end-to-end autonomous driving systems often leverage large language models to generate driving decisions directly based on their understanding of the current scene. However, such systems introduce multiple risks in real-world driving scenarios. To evaluate whether VLMs are truly viable for autonomous driving, we introduce RoboDriveBench, the first robustness benchmark focused on end-to-end trajectory prediction tasks. This benchmark systematically evaluates two critical categories of real-world challenges for VLM-based end-to-end autonomous driving systems through 11 simulated scenarios encompassing various corruption types, including 6 scenarios of sensor corruption caused by environmental variations, along with 5 cases of prompt corruption resulting from human intervention and data transmission failures. Each corruption type includes 250 unique driving scenarios and 5,689 frames, resulting in 64,559 total trajectory prediction cases per evaluation. To overcome these real-world challenges, we propose a novel VLM-based autonomous driving framework called RoboDriveVLM, which enhances robustness by mapping more multimodal data-e.g., lidar and radar-into a unified latent space. Furthermore, we introduce a new Test-Time Adaptation (TTA) method based on cross-modal knowledge distillation to improve the robustness of VLM-based autonomous driving systems. Through extensive experiments, our work highlights the limitations of current VLM-based end-to-end autonomous driving systems and provides a more reliable solution for real-world deployment. Source code and datasets will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。