arXiv:2606.09142cs.CVcs.AI2026-06

用视觉语言模型从第一视角视频预测行人过街意图,准确率提升14.5%。

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models

论文配图:Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models
图 1 · 摘自论文原文
  • 将过街意图识别转为闭合式视觉问答,利用VLM模型进行推理
  • 微调后模型相比基线提升14.5%准确率,达新最佳水平
  • 结合眼动与自车运动信息,显著增强预测性能,适合自动驾驶研究者

第一人称视觉提供了人类感知与决策的直接视角,但在交通安全性预测中的潜力尚未充分挖掘。本文研究从短时第一人称视频片段中解码行人过街意图。通过将任务定义为闭合式视觉问答(VQA),并利用视觉语言模型(VLM)预测行人意图。我们首先在零样本设置下评估三种先进VLM,发现其表现优于随机猜测但高层交通推理能力有限。为此,我们采用参数高效微调适配目标任务。结果表明,微调模型显著优于零样本版本,较专用Transformer基线提升9%准确率。进一步引入自车运动、车辆运动及眼动等上下文线索,性能再获提升。特别是基于眼动与自车运动引导的微调Qwen3-VL-2B模型,相较基线实现14.5%准确率提升,建立第一人称行人意图解码新基准。

原文摘要 · Abstract (English)

Egocentric vision offers a first-person view of human perception and decision making, yet its potential for traffic-safety prediction remains underexplored. In this work, we study the decoding of pedestrian crossing intentions from short egocentric video clips. We approach this by formulating the task as a closed-ended visual question answering (VQA) problem and leveraging vision language models (VLMs) to predict the pedestrians' intent. We first benchmark three families of state-of-the-art VLMs in a zero-shot setting, finding that they achieve moderate gains over random guessing but exhibit limited higher-level traffic reasoning. Motivated by these findings, we further adapt VLMs to the target task using parameter-efficient fine-tuning. Our results show that the fine-tuned models substantially outperform their zero-shot counterparts and achieve a 9\% accuracy improvement over a specialized transformer-based baseline. Finally, we demonstrate that incorporating additional contextual cues, including ego motion, vehicle motion, and eye gaze, further improves predictive performance. In particular, the fine-tuned Qwen3-VL-2B model guided by eye gaze and ego motion achieves a 14.5% accuracy improvement over the transformer baseline, establishing a new state of the art for egocentric pedestrian intent decoding.

视觉语言模型行为预测自动驾驶第一人称视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。