arXiv:2606.30807cs.ROcs.CR2026-06

攻击自动驾驶生成模型的评分头,让其选中危险路径。

Off the Rails: Hijacking the Scoring Head in Generative End-to-End Driving Planners with Safety-Violating Adversarial Perturbations

论文配图:Off the Rails: Hijacking the Scoring Head in Generative End-to-End Driving Planners with Safety-Violating Adversarial Perturbations
图 1 · 摘自论文原文
  • 通过扰动评分头诱导模型选择危险轨迹。
  • 使安全路径得分下降39%至80%,碰撞率最高达50%。
  • 揭示评分头是关键攻击面,适合安全研究者关注。

生成模型近年来在端到端自动驾驶中迅速普及,基于扩散去噪和基于词表检索成为主流轨迹解码范式。尽管架构多样,当前生成式自动驾驶规划器共享同一推理模式:固定候选轨迹集由一个或多个基于鸟瞰图(BEV)特征的可学习评分头打分,得分最高的候选轨迹被选为最终输出。在此设计下,评分头是感知与运动指令间唯一的屏障,且不同候选之间的决策边界常很微弱。我们提出 extsc{Derail},一种针对评分头的对抗性攻击框架。在多种生成式规划器上评估, extsc{Derail} 能将轨迹选择从安全转向危险,使得分下降39%–80%,碰撞率最高达50%,显著优于通用损失最大化和特征差异攻击。分析表明,违反安全的目标决定了攻击效果,而评分头的推理模式本身即为重复出现的攻击面,值得专门防御考虑。

原文摘要 · Abstract (English)

Generative models have recently seen rapid adoption in End-to-End (E2E) autonomous driving (AD), with diffusion-based denoising and vocabulary-based retrieval becoming the dominant trajectory-decoding paradigms. Despite their architectural diversity, current generative AD planners share a common inference pattern: a fixed set of candidate trajectories (anchors, vocabulary entries, or proposal queries) is scored by one or more learned heads conditioned on the Bird's-Eye-View (BEV) features, and the highest-scored candidate is returned as the final trajectory. Under this design, the scoring head is the only barrier between perception and the motion command, and its decision margins between competing candidates are often small. We introduce \textsc{Derail}, an adversarial framework that exploits this scoring-head attack surface. Evaluated on various generative planners, \textsc{Derail} flips the trajectory selection from a safe to an unsafe candidate, with score drops of $39$--$80\%$ and collision rates of up to $50\%$, consistently outperforming generic loss-maximization and feature-divergence attacks. Our analysis suggests that safety-violating objectives govern attack effectiveness against generative AD planners, and that the scoring-head inference pattern itself is a recurring attack surface worth explicit defensive consideration.

自动驾驶对抗攻击生成模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。