arXiv:2607.06328cs.AIcs.CV2026-07

用可解释性模块分析端到端自动驾驶模型,发现并修正错误决策。

Driving the Wrong Way: Leveraging Interpretability in End2End Autonomous Driving Models

论文配图:Driving the Wrong Way: Leveraging Interpretability in End2End Autonomous Driving Models
图 1 · 摘自论文原文
  • 引入无监督字典学习分解驾驶行为为语义概念
  • 通过概念干预提升驾驶性能,验证因果影响
  • 适合关注自动驾驶安全与模型透明度的研究者

端到端自动驾驶模型因复杂性和黑箱特性,可能学习到错误行为。本文在先进驾驶模型中集成无监督字典学习作为后处理可解释性模块,将驾驶行为分解为具有语义意义的概念,并证明这些概念对模型决策的因果影响。提出分步框架,从模型中提取并解释有意义的概念,将其与多维度输出关联,揭示未来轨迹预测的决策逻辑。针对概念层面的定向干预可有效修正驾驶决策,显著提升整体性能。结果表明,可解释性有助于降低模型不透明性,发现错误行为,并实现精准修正,最终增强模型表现。

原文摘要 · Abstract (English)

The increasing adoption of end-to-end learning for autonomous driving introduces increased model complexity and opacity, raising the risk of learning undesired or erroneous behavior. In this work, we integrate unsupervised dictionary learning as a post hoc interpretability module within state-of-the-art driving models to decompose driving behavior into semantically meaningful concepts while demonstrating their causal influence on the model's driving decisions. We propose a stepwise framework for extracting and interpreting meaningful concepts from the end-to-end model and connecting them to the multifaceted model outputs, thereby revealing the underlying decision-making logic for the prediction of future trajectories. Furthermore, targeted interventions at the concept level allow us to manipulate and correct driving decisions, resulting in measurable improvements in overall driving performance. We thus demonstrate how interpretability can effectively be used to reduce model opacity, uncover erroneous behavior, and enable targeted mitigation, ultimately boosting model performance.

自动驾驶可解释性端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。