arXiv:2510.09200cs.CVcs.AI2025-10ICCV

让自动驾驶预测司机意图更安全可解释。

Towards Safer and Understandable Driver Intention Prediction

  • 用视频概念瓶颈模型生成时空连贯的决策解释。
  • 基于新数据集验证,Transformer比CNN更具可解释性。
  • 适合关注自动驾驶可解释性的研究者和工程师。

自动驾驶系统日益依赖深度学习处理复杂任务,但人机交互中决策过程的可解释性对保障行车安全至关重要。为此,本文提出在变道等行为发生前进行可解释的司机意图预测(DIP),并构建了多模态、以车辆为中心的视频数据集DAAD-X,提供来自驾驶员眼动和车辆视角的分层高层文本解释。提出视频概念瓶颈模型(VCBM),能自动生成时空一致的解释,无需事后分析。在DAAD-X上的实验证明,基于Transformer的模型比传统CNN模型具有更强可解释性;同时引入多标签t-SNE可视化方法,展示多个解释间的解耦与因果关联。相关数据、代码和模型已公开。

原文摘要 · Abstract (English)

Autonomous driving (AD) systems are becoming increasingly capable of handling complex tasks, mainly due to recent advances in deep learning and AI. As interactions between autonomous systems and humans increase, the interpretability of decision-making processes in driving systems becomes increasingly crucial for ensuring safe driving operations. Successful human-machine interaction requires understanding the underlying representations of the environment and the driving task, which remains a significant challenge in deep learning-based systems. To address this, we introduce the task of interpretability in maneuver prediction before they occur for driver safety, i.e., driver intent prediction (DIP), which plays a critical role in AD systems. To foster research in interpretable DIP, we curate the eXplainable Driving Action Anticipation Dataset (DAAD-X), a new multimodal, ego-centric video dataset to provide hierarchical, high-level textual explanations as causal reasoning for the driver's decisions. These explanations are derived from both the driver's eye-gaze and the ego-vehicle's perspective. Next, we propose Video Concept Bottleneck Model (VCBM), a framework that generates spatio-temporally coherent explanations inherently, without relying on post-hoc techniques. Finally, through extensive evaluations of the proposed VCBM on the DAAD-X dataset, we demonstrate that transformer-based models exhibit greater interpretability than conventional CNN-based models. Additionally, we introduce a multilabel t-SNE visualization technique to illustrate the disentanglement and causal correlation among multiple explanations. Our data, code and models are available at: https://mukil07.github.io/VCBM.github.io/

自动驾驶意图预测可解释性视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。