构建首个聚焦长尾场景推理的自动驾驶数据集,助力系统做出更安全、可解释的决策。
nuReasoning: A Reasoning-Centric Dataset and Benchmark for Long-Tail Autonomous Driving

- 构建2万段真实道路视频,含多传感器数据与人类标注的三类推理标签。
- 在推理监督下训练模型,显著提升问答准确率与规划性能,即使推理输出不参与推断。
- 适合研究自动驾驶中常识推理、决策可解释性与长尾场景鲁棒性的团队使用。
推理对自动驾驶在长尾场景中的表现至关重要,车辆需运用常识知识、理解空间关系、推断行为交互并作出安全决策。然而现有数据集和基准主要聚焦感知、预测或规划,对长尾驾驶场景中的推理监督支持有限。我们提出nuReasoning,一个大规模真实世界数据集与基准,专为推理驱动的自动驾驶设计。数据集包含20,000个片段,每段20秒,覆盖多个城市,提供同步的多摄像头图像、激光雷达数据、高精地图、目标标注及人工验证的推理标注,涵盖空间推理、决策推理和反事实推理三类。不同于以往以视觉问答为主的基准,nuReasoning同时支持推理评估与规划评估,可直接研究推理监督对驾驶性能的影响。实验表明,在nuReasoning上微调视觉语言模型(VLM)显著提升驾驶相关问答能力;将推理监督融入视觉语言动作模型(VLA)训练,即使推理输出在推理时禁用,仍能改善规划性能。这些结果确立了nuReasoning作为评估与提升真实长尾场景下鲁棒、可解释自动驾驶系统的基石。
原文摘要 · Abstract (English)
Reasoning is essential for autonomous driving (AD) in long-tail scenarios, where vehicles must apply commonsense knowledge, understand spatial relations, infer agent interactions, and make safe decisions. However, existing AD datasets and benchmarks mainly target perception, prediction, or planning, and provide limited supervision for reasoning over realistic long-tail driving scenes. We introduce nuReasoning, a large-scale real-world dataset and benchmark for reasoning-centric AD. Following the lineage of nuScenes and nuPlan, nuReasoning advances real-world AD datasets and benchmarks toward reasoning in long-tail driving scenarios. The dataset contains 20,000 clips, each 20 seconds long, collected across multiple cities, with synchronized multi-camera images, LiDAR data, HD maps, object annotations, and human-verified reasoning annotations spanning Spatial Reasoning, Decision Reasoning, and Counterfactual Reasoning. Unlike prior datasets that focus primarily on visual question answering, nuReasoning supports both reasoning evaluation and planning evaluation, enabling a direct study of how reasoning supervision affects driving performance. Experiments show that fine-tuning VLMs on nuReasoning substantially improves driving-specific question answering, while incorporating reasoning supervision into VLA training improves planning performance even when textual reasoning outputs are disabled at inference time. These results establish nuReasoning as a foundation for evaluating and improving robust, interpretable, reasoning-driven AD systems in realistic long-tail settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。