不训练就能预测行人过街意图,靠连续视频+环境信息
Seeing Beyond Frames: Zero-Shot Pedestrian Intention Prediction with Raw Temporal Video and Multimodal Cues
- 用连续视频片段+结构化元数据做零样本推理
- 零训练下准确率达73%,比GPT-4V高18%
- 适合需要快速适配新场景的自动驾驶系统
行人意图预测对复杂城市环境中的自动驾驶至关重要。传统方法依赖帧序列的监督学习,需大量重训才能适应新场景。本文提出BF-PIP(Beyond Frames Pedestrian Intention Prediction),基于Gemini 2.5 Pro构建的零样本方法,直接从含结构化JAAD元数据的短时连续视频片段中推断过街意图。与基于离散帧的GPT-4V方法不同,BF-PIP处理不间断时间片段,并通过专用多模态提示融合边界框标注与自车速度信息。无需额外训练,其预测准确率达到73%,较GPT-4V基线提升18%。结果表明,结合时序视频输入与上下文线索可增强时空感知,在模糊条件下改善意图推断。该方法为智能交通系统提供了无需重训的敏捷感知模块。
原文摘要 · Abstract (English)
Pedestrian intention prediction is essential for autonomous driving in complex urban environments. Conventional approaches depend on supervised learning over frame sequences and require extensive retraining to adapt to new scenarios. Here, we introduce BF-PIP (Beyond Frames Pedestrian Intention Prediction), a zero-shot approach built upon Gemini 2.5 Pro. It infers crossing intentions directly from short, continuous video clips enriched with structured JAAD metadata. In contrast to GPT-4V based methods that operate on discrete frames, BF-PIP processes uninterrupted temporal clips. It also incorporates bounding-box annotations and ego-vehicle speed via specialized multimodal prompts. Without any additional training, BF-PIP achieves 73% prediction accuracy, outperforming a GPT-4V baseline by 18 %. These findings illustrate that combining temporal video inputs with contextual cues enhances spatiotemporal perception and improves intent inference under ambiguous conditions. This approach paves the way for agile, retraining-free perception module in intelligent transportation system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。