arXiv:2506.17590cs.CVcs.AI2025-06ICCV被引 8

构建驾驶场景下行人意图与风险推理的细粒度评测基准

DRAMA-X: A Fine-grained Intent Prediction and Risk Reasoning Benchmark For Driving

  • 基于自动化标注构建多类别意图分类体系
  • 9类方向意图+风险评分,支持四类决策任务评估
  • 适合自动驾驶安全决策研究者参考

理解行人、骑行者等弱势道路使用者的短期运动行为对自动驾驶安全至关重要,尤其在城市环境中存在模糊或高风险行为时。尽管视觉语言模型(VLMs)已实现开放词汇感知,但其在细粒度意图推理中的应用仍待探索。现有缺乏针对高危情境下多类别意图预测的评测基准。为此,我们提出DRAMA-X,基于DRAMA数据集通过自动化标注流程构建。该基准包含5,686个事故高发帧,标注了目标边界框、九类方向意图、二值风险评分、专家生成的自车动作建议及运动描述摘要。这些标注支持对四个核心自动驾驶决策任务的结构化评估:目标检测、意图预测、风险评估和动作建议。作为基线,我们提出SGG-Intent,一种轻量级、无需训练的框架,模拟自车推理流程:利用VLM驱动检测器生成场景图,推断意图,评估风险,并通过大语言模型驱动的组合推理阶段推荐动作。我们评估多种近期VLMs在所有四项任务上的表现。实验表明,基于场景图的推理能提升意图预测与风险评估效果,尤其是在显式建模上下文线索时。

原文摘要 · Abstract (English)

Understanding the short-term motion of vulnerable road users (VRUs) like pedestrians and cyclists is critical for safe autonomous driving, especially in urban scenarios with ambiguous or high-risk behaviors. While vision-language models (VLMs) have enabled open-vocabulary perception, their utility for fine-grained intent reasoning remains underexplored. Notably, no existing benchmark evaluates multi-class intent prediction in safety-critical situations, To address this gap, we introduce DRAMA-X, a fine-grained benchmark constructed from the DRAMA dataset via an automated annotation pipeline. DRAMA-X contains 5,686 accident-prone frames labeled with object bounding boxes, a nine-class directional intent taxonomy, binary risk scores, expert-generated action suggestions for the ego vehicle, and descriptive motion summaries. These annotations enable a structured evaluation of four interrelated tasks central to autonomous decision-making: object detection, intent prediction, risk assessment, and action suggestion. As a reference baseline, we propose SGG-Intent, a lightweight, training-free framework that mirrors the ego vehicle's reasoning pipeline. It sequentially generates a scene graph from visual input using VLM-backed detectors, infers intent, assesses risk, and recommends an action using a compositional reasoning stage powered by a large language model. We evaluate a range of recent VLMs, comparing performance across all four DRAMA-X tasks. Our experiments demonstrate that scene-graph-based reasoning enhances intent prediction and risk assessment, especially when contextual cues are explicitly modeled.

自动驾驶意图预测风险评估视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。