arXiv:2605.24562cs.CVcs.AI2026-05被引 1

用问答形式构建行人行为预测数据集,提升模型理解与解释能力。

PEDESTRIANQA: A Benchmark for Vision-Language Models on Pedestrian Intention and Trajectory Prediction

论文配图:PEDESTRIANQA: A Benchmark for Vision-Language Models on Pedestrian Intention and Trajectory Prediction
图 1 · 摘自论文原文
  • 将行人意图与轨迹预测转化为带结构化推理的问答任务
  • 在多个数据集上显著提升预测准确率与解释质量
  • 适合自动驾驶安全系统研发及多模态模型研究者

行人意图与轨迹预测对自动驾驶系统的安全部署至关重要,直接影响复杂交通环境下的导航决策。近年来,大型视觉语言模型(VLMs)通过融合强大的视觉理解与灵活的自然语言推理,为该任务提供了新范式。本文提出PedestrianQA,一个大规模基于视频的数据集,将行人意图与轨迹预测转化为带有结构化推理的问答任务。该数据集以自然语言形式丰富标注行人序列,使VLMs能够从视觉动态、上下文线索及交通参与者间交互中学习,并生成简洁的预测解释,无需针对每项任务设计专用架构。在PIE、JAAD、TITAN和IDD-PeD上的实证评估表明,微调先进VLMs可显著提升意图分类准确率、轨迹预测精度以及解释性,验证了VLMs作为统一且可解释框架在安全关键行人行为建模中的巨大潜力。

原文摘要 · Abstract (English)

Pedestrian intention and trajectory prediction are critical for the safe deployment of autonomous driving systems, directly influencing navigation decisions in complex traffic environments. Recent advances in large vision-language models offer a powerful new paradigm for these tasks by combining high-capacity visual understanding with flexible natural language reasoning. In this work, we introduce PedestrianQA, a large-scale video-based dataset that formulates pedestrian intention and trajectory prediction as question-answering tasks augmented with structured rationales. PedestrianQA expresses richly annotated pedestrian sequences, in natural language, enabling VLMs to learn from visual dynamics, contextual cues, and interactions among traffic agents while generating concise explanations of their predictions without needing specialized architectures tailored for each task. Empirical evaluations across PIE, JAAD, TITAN, and IDD-PeD show that finetuning state-of-the-art VLMs on PedestrianQA significantly improves intention classification, trajectory forecasting accuracy, and the quality of explanatory rationales, demonstrating the strong potential of VLMs as a unified and explainable framework for safety-critical pedestrian behavior modeling.

行人预测视觉语言模型自动驾驶可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。