arXiv:2608.04896cs.AIcs.CV2026-08

共享模拟轨迹会误导自动驾驶评估,导致无效策略得分高于人类。

When Shared Rollouts Fail in Defensive Driving Evaluation: A NAVSIM Score Basis Audit

论文配图:When Shared Rollouts Fail in Defensive Driving Evaluation: A NAVSIM Score Basis Audit
图 1 · 摘自论文原文
  • 通过共享轨迹的模拟评估,误将参考失败归因于智能体。
  • 在12,146个测试样本中,盲探策略得分超过人类与先进模型。
  • 提出审计协议,强调数值稳定与透明性,适用于安全评估研究者。

防御性驾驶评分仅在能区分观察环境与不观察环境的策略时才有效。重模拟基准可能采用参考条件宽容机制,即当记录的人类参考未通过合规通道时,智能体可获得奖励。当智能体与参考共享不稳定的轨迹变换时,该规则会将共同的参考失败传播为广泛的合规信用。我们在NAVSIM v2.2原始场景单阶段评分中审计了这一风险。在受控的数值后端条件下,路线无关的忽略全部探测器和路线感知但对象盲视的探测器,在完整的12,146个token导航测试集中得分高于人类回放和PDM-Closed。按照公开规范重新安装后,在固定32个token诊断集上重现了轨迹发散。同一源依赖栈控制与精确输入诊断隔离了共享速度重拟合中的依赖敏感数值行为。在450个token控制池中,仅替换求解器即可消除轨迹发散,恢复盲视最后的排序,同时保持宽容机制启用。因此,数值不稳定是直接诱因。参考条件宽容将由此产生的共享参考失败传播为合规信用。我们贡献了一套审计协议,要求披露评分基础与依赖栈、使用盲探、覆盖报告及轨迹稳定性测试,方可用于防御性驾驶主张。

原文摘要 · Abstract (English)

Defensive driving scores are useful only when they preserve distinctions between policies that observe surrounding actors and those that do not. Re-simulation benchmarks may use reference-conditioned forgiveness, under which an agent receives credit when the logged human reference fails a compliance channel. When agent and reference share an unstable rollout transformation, this rule can propagate shared reference failures into broad compliance credit. We audit this risk in NAVSIM v2.2 original scene single-stage scoring. Under the affected documented-stack condition on the audited numerical backend, the route-blind Ignore-All probe and a route-aware actor-blind probe outrank human replay and PDM-Closed over the complete 12,146-token navtest split. A fresh installation following the public specification reproduces rollout divergence on a fixed 32-token diagnostic set. A same-source dependency stack control and an exact-input diagnostic isolate dependency-sensitive numerical behavior in the shared velocity refit. On a 450-token control pool, replacing only the solver eliminates rollout divergence and restores blind-last ordering while keeping forgiveness enabled. Thus, the numerical instability is the direct trigger. Reference-conditioned forgiveness propagates the resulting shared reference failures into compliance credit. We contribute an audit protocol requiring score basis and stack disclosure, blind probes, overwrite reporting, and rollout stability tests before using such scores for defensive driving claims.

自动驾驶评估审计数值稳定

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。