arXiv:2507.04141cs.CVcs.AI2025-07被引 9

用视觉语言模型预测行人过街意图,提升自动驾驶理解能力。

Pedestrian Intention Prediction via Vision-Language Foundation Models

  • 通过分层提示模板融合视觉、物理线索和车辆动态信息
  • 在三数据集上准确率最高提升19.8%,自动提示优化再增12.5%
  • 适合关注多模态感知与自动驾驶决策的研究者

行人过街意图预测是自动驾驶的关键功能。传统基于视觉的方法在泛化能力、上下文理解与因果推理方面存在不足。本文探索视觉语言基础模型(VLFMs)在行人过街意图预测中的潜力,通过分层提示模板整合多模态数据。方法融合视觉帧、物理线索观测及自车动态信息,构建系统化优化的提示以引导VLFMs进行意图预测。在JAAD、PIE和FU-PIP三个常用数据集上进行实验,结果表明:引入车辆速度、其时序变化及时间感知提示可将预测准确率提升至19.8%;通过自动提示工程框架生成的优化提示进一步带来12.5%的准确率增益。这些结果凸显了VLFMs相较于传统视觉模型的优越性能,为自动驾驶应用提供了更强的泛化能力和上下文理解力。

原文摘要 · Abstract (English)

Prediction of pedestrian crossing intention is a critical function in autonomous vehicles. Conventional vision-based methods of crossing intention prediction often struggle with generalizability, context understanding, and causal reasoning. This study explores the potential of vision-language foundation models (VLFMs) for predicting pedestrian crossing intentions by integrating multimodal data through hierarchical prompt templates. The methodology incorporates contextual information, including visual frames, physical cues observations, and ego-vehicle dynamics, into systematically refined prompts to guide VLFMs effectively in intention prediction. Experiments were conducted on three common datasets-JAAD, PIE, and FU-PIP. Results demonstrate that incorporating vehicle speed, its variations over time, and time-conscious prompts significantly enhances the prediction accuracy up to 19.8%. Additionally, optimised prompts generated via an automatic prompt engineering framework yielded 12.5% further accuracy gains. These findings highlight the superior performance of VLFMs compared to conventional vision-based models, offering enhanced generalisation and contextual understanding for autonomous driving applications.

行人预测视觉语言模型自动驾驶多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。