arXiv:2509.25944cs.AI2025-09被引 6

构建首个面向自动驾驶风险评估的时序视觉问答数据集,支持精细化行为推理。

NuRisk: A Visual Question Answering Dataset for Agent-Level Risk Assessment in Autonomous Driving

  • 基于鸟瞰图序列图像与真实数据,构建2.9千场景的多智能体风险标注数据集
  • 现有模型最高准确率仅33%,且延迟高,难以实现时序推理
  • 微调后7B模型提升至41%准确率,降低75%延迟,验证显式时序推理能力

自动驾驶中的风险理解不仅需要感知与预测,还需对智能体行为和环境上下文进行高层推理。当前基于视觉语言模型(VLM)的方法主要依赖静态图像,提供定性判断,缺乏捕捉风险随时间演变所需的时空推理能力。为此,我们提出NuRisk,一个全面的视觉问答(VQA)数据集,包含2.9K场景和110万条智能体级样本,基于nuScenes和Waymo的真实数据,并补充了CommonRoad模拟器中的安全关键场景。该数据集提供基于鸟瞰图(BEV)的序列图像及定量的智能体级风险标注,支持时空推理。我们在多种提示技术下对知名VLM进行基准测试,发现其无法进行显式的时空推理,最高准确率仅为33%,且延迟较高。为解决此问题,我们微调的7B VLM模型将准确率提升至41%,延迟降低75%,展现出显式时空推理能力,而该能力在现有专有模型中未见。尽管取得显著进展,但准确率仍较低,凸显任务难度,确立了NuRisk作为推动自动驾驶时空推理的关键基准。

原文摘要 · Abstract (English)

Understanding risk in autonomous driving requires not only perception and prediction, but also high-level reasoning about agent behavior and context. Current Vision Language Model (VLM)-based methods primarily ground agents in static images and provide qualitative judgments, lacking the spatio-temporal reasoning needed to capture how risks evolve over time. To address this gap, we propose NuRisk, a comprehensive Visual Question Answering (VQA) dataset comprising 2.9K scenarios and 1.1M agent-level samples, built on real-world data from nuScenes and Waymo, completed with safety-critical scenarios from the CommonRoad simulator. The dataset provides Bird's-eye view (BEV) based sequential images with quantitative, agent-level risk annotations, enabling spatio-temporal reasoning. We benchmark well-known VLMs across different prompting techniques and find that they fail to perform explicit spatio-temporal reasoning, resulting in a peak accuracy of 33% at high latency. To address these shortcomings, our fine-tuned 7B VLM agent improves accuracy to 41% and reduces latency by 75%, demonstrating explicit spatio-temporal reasoning capabilities that proprietary models lacked. While this represents a significant step forward, the modest accuracy underscores the profound challenge of the task, establishing NuRisk as a critical benchmark for advancing spatio-temporal reasoning in autonomous driving. More information can be found at https://github.com/TUM-AVS/NuRisk.

自动驾驶风险评估视觉问答时空推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。