构建真实驾驶场景的多模态数据集,让自动驾驶系统学会用人类能理解的方式解释决策。
NARRATE: A Multimodal Real-World Australian Driving Dataset for Human-Centred Explanations in Automated Driving

- 从35名驾驶员真实道路驾驶中采集2050个事件,同步记录视觉、定位、运动与激光雷达数据。
- 每条事件配有车内或事后自由文本解释,并标注感知、理解、预测三阶段情境意识。
- 可用于训练更贴近人类认知的自动驾驶解释模型,尤其适合关注人因与可解释性的研究者。
自动驾驶车辆需以乘客可理解、可监控、可信任的方式解释其决策。现有语言标注驾驶数据集多为观察者事后撰写、仿真生成或基于传感器输入,而非来自实际执行动作的驾驶员。我们提出NARRATE,一个面向真实澳大利亚道路的多模态驾驶数据集,包含35名经验丰富的驾驶员与教练在公共道路上完成的2050个标注事件。每个事件均同步采集视觉、定位、运动与LiDAR数据,并配以车内或事后自由文本解释。数据集提供动作标签、涵盖六个高层级和三十二个细粒度类别的场景上下文标签,以及对驾驶员解释中感知、理解、预测三个阶段的情境意识(SA)分段标注。四个基准任务(情境意识、场景上下文、驾驶行为分类与解释生成)表明,驾驶员语言中的结构可被学习,但细粒度上下文识别与解释生成仍具挑战。NARRATE为构建更以人为中心、更具领域感知能力的自动驾驶解释模型开辟了路径。
原文摘要 · Abstract (English)
Automated vehicles must explain their decisions in ways that passengers can understand, monitor, and trust. Existing language-annotated driving datasets are mostly observer-written, post-hoc, simulation-based, or generated from sensor inputs, rather than elicited from the driver performing the action. We introduce NARRATE, a multimodal real-world Australian driving dataset comprising 2,050 annotated events from 35 experienced drivers and driving instructors on public roads. Each event is grounded in synchronised visual, localisation, motion, and LiDAR streams and paired with in-vehicle and/or post-drive free-text explanations. NARRATE provides action labels, scenario-context labels spanning six high-level and 32 fine-grained categories, and span-level Situational Awareness (SA) annotations over driver explanations for Perception, Comprehension and Projection. Four benchmark tasks (SA, scenario-context, driver-action classification, and explanation generation) show that this structure is learnable from driver language, while fine-grained context recognition and explanation generation remain challenging. NARRATE paves a path towards more human-centred and domain-aware explanation models for automated driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。