arXiv:2606.31106cs.ROcs.AI2026-06

通过探针分析发现,自动驾驶模型在危机时刻缺乏对周围车辆的预测能力。

What Probing Reveals about Autonomous Driving: Linking Internal Prediction Errors to Ego Planning

论文配图:What Probing Reveals about Autonomous Driving: Linking Internal Prediction Errors to Ego Planning
图 1 · 摘自论文原文
  • 用线性探针和扰动实验检测模型内部的预测与规划信号。
  • 大模型虽闭环表现好,但碰撞前难以及时预测周围车辆位置。
  • 修正错误预测可显著提升自身轨迹安全性,适合安全评估研究者。

大规模数据集和快速模拟器推动了驾驶策略的进步,使其在常规场景中表现安全可靠,但良好性能可能掩盖推理缺陷和不安全启发式。闭环模拟器的综合评分无法揭示策略本质,难以判断其是否真正预测周边车辆运动、如何生成自身路径,还是仅依赖在常规场景中偶然奏效的脆弱规则。为理解驾驶策略的局限性,我们聚焦于探测预测(即周围车辆下一步位置)与规划(即如何生成安全轨迹)能力。这两项能力反映有效驾驶策略的核心要求,我们以此评估基于行为克隆和模拟强化学习的策略质量。通过分析规模影响,考察更大数据集和更长训练是否带来更强预测与规划能力,而非仅改善行为启发式。在线性探针与针对性扰动实验中,追踪预测与规划信号的出现、饱和或失效。结果显示,尽管闭环性能优异,模型在近碰撞事件中仍常无法及时形成对周围车辆的预测,暴露出规划所需预测信号的不足。因果干预表明,纠正错误预测可显著改善自身轨迹的安全性。

原文摘要 · Abstract (English)

Large-scale datasets and fast simulators have enabled improvements in driving policies that appear safe and robust, yet strong performance in nominal scenarios can still mask flawed reasoning and unsafe heuristics. Summary scores from closed-loop simulators do not give significant insight into the policy, making it difficult to determine whether they truly predict the motion of surrounding vehicles, how the ego vehicle generates future plans, or whether they merely rely on brittle heuristics that happen to succeed in nominal scenarios. To better understand the limits and weaknesses of driving policies, we focus on probing for forms of prediction, i.e., where surrounding vehicles will move next, and planning, i.e., understanding how to generate safe trajectories. We focus on these two capabilities because they reflect behaviors expected of effective driving policies, and use their presence or absence to assess policy quality across data-driven behavior cloning and simulation-driven reinforcement learning policies. To evaluate the presence of these capabilities, we investigate them as a function of scale, asking whether the closed-loop gains from larger datasets and longer simulation training reflect stronger prediction and planning or merely better behavioral heuristics. We use linear probing and targeted perturbations in both imitation learning and reinforcement learning models to track when these internal signals emerge, plateau, or fail. Despite good closed-loop performance, policies often fail to form timely surrounding-vehicle predictions during near-collision events, revealing a limitation in the predictive signals available for ego planning. Finally, causal intervention shows that correcting mistaken predictions improves ego planning toward safer trajectories.

自动驾驶模型探针安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。