构建驾驶场景下行人意图与风险推理的细粒度评测基准
DRAMA-X: A Fine-grained Intent Prediction and Risk Reasoning Benchmark For Driving
- 基于自动化标注构建多类别意图分类体系
- 9类方向意图+风险评分,支持四类决策任务评估
- 适合自动驾驶安全决策研究者参考
理解行人、骑行者等弱势道路使用者的短期运动行为对自动驾驶安全至关重要,尤其在城市环境中存在模糊或高风险行为时。尽管视觉语言模型(VLMs)已实现开放词汇感知,但其在细粒度意图推理中的应用仍待探索。现有缺乏针对高危情境下多类别意图预测的评测基准。为此,我们提出DRAMA-X,基于DRAMA数据集通过自动化标注流程构建。该基准包含5,686个事故高发帧,标注了目标边界框、九类方向意图、二值风险评分、专家生成的自车动作建议及运动描述摘要。这些标注支持对四个核心自动驾驶决策任务的结构化评估:目标检测、意图预测、风险评估和动作建议。作为基线,我们提出SGG-Intent,一种轻量级、无需训练的框架,模拟自车推理流程:利用VLM驱动检测器生成场景图,推断意图,评估风险,并通过大语言模型驱动的组合推理阶段推荐动作。我们评估多种近期VLMs在所有四项任务上的表现。实验表明,基于场景图的推理能提升意图预测与风险评估效果,尤其是在显式建模上下文线索时。
原文摘要 · Abstract (English)
Understanding the short-term motion of vulnerable road users (VRUs) like pedestrians and cyclists is critical for safe autonomous driving, especially in urban scenarios with ambiguous or high-risk behaviors. While vision-language models (VLMs) have enabled open-vocabulary perception, their utility for fine-grained intent reasoning remains underexplored. Notably, no existing benchmark evaluates multi-class intent prediction in safety-critical situations, To address this gap, we introduce DRAMA-X, a fine-grained benchmark constructed from the DRAMA dataset via an automated annotation pipeline. DRAMA-X contains 5,686 accident-prone frames labeled with object bounding boxes, a nine-class directional intent taxonomy, binary risk scores, expert-generated action suggestions for the ego vehicle, and descriptive motion summaries. These annotations enable a structured evaluation of four interrelated tasks central to autonomous decision-making: object detection, intent prediction, risk assessment, and action suggestion. As a reference baseline, we propose SGG-Intent, a lightweight, training-free framework that mirrors the ego vehicle's reasoning pipeline. It sequentially generates a scene graph from visual input using VLM-backed detectors, infers intent, assesses risk, and recommends an action using a compositional reasoning stage powered by a large language model. We evaluate a range of recent VLMs, comparing performance across all four DRAMA-X tasks. Our experiments demonstrate that scene-graph-based reasoning enhances intent prediction and risk assessment, especially when contextual cues are explicitly modeled.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。