arXiv:2409.06702cs.CVcs.AI2024-09CoRL被引 35

让自动驾驶语言解释与系统输出全程对齐,提升可解释性。

Hint-AD: Holistically Aligned Interpretability in End-to-End Autonomous Driving

  • 将自然语言与感知-预测-规划全流程输出对齐生成
  • 在驾驶解释等任务上达到当前最优性能
  • 适合关注自动驾驶可解释性的研究者使用

端到端自动驾驶系统面临可解释性挑战,影响人机信任。现有工作多聚焦于陈述式解释,语言与系统中间输出无关联。本文提出 Hint-AD,一个集成的自动驾驶-语言系统,生成与感知-预测-规划全过程输出一致的语言解释。通过融合中间特征并引入整体令牌混合子网络,实现有效特征适配,在驾驶解释、3D密集描述和指令预测等任务上取得当前最佳表现。为促进 nuScenes 上驾驶解释研究,我们构建了人工标注数据集 Nu-X。代码、数据集与模型将公开。

原文摘要 · Abstract (English)

End-to-end architectures in autonomous driving (AD) face a significant challenge in interpretability, impeding human-AI trust. Human-friendly natural language has been explored for tasks such as driving explanation and 3D captioning. However, previous works primarily focused on the paradigm of declarative interpretability, where the natural language interpretations are not grounded in the intermediate outputs of AD systems, making the interpretations only declarative. In contrast, aligned interpretability establishes a connection between language and the intermediate outputs of AD systems. Here we introduce Hint-AD, an integrated AD-language system that generates language aligned with the holistic perception-prediction-planning outputs of the AD model. By incorporating the intermediate outputs and a holistic token mixer sub-network for effective feature adaptation, Hint-AD achieves desirable accuracy, achieving state-of-the-art results in driving language tasks including driving explanation, 3D dense captioning, and command prediction. To facilitate further study on driving explanation task on nuScenes, we also introduce a human-labeled dataset, Nu-X. Codes, dataset, and models will be publicly available.

自动驾驶可解释性语言生成多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。