arXiv:2511.14120cs.CVcs.AI2025-11被引 1

用多视角+语言模型分析行人事故全过程,让AI懂行为阶段并提预防建议。

Multi-view Phase-aware Pedestrian-Vehicle Incident Reasoning Framework with Vision-Language Models

  • 分四阶段处理多视角视频,按行人认知阶段拆解事故过程。
  • 行为阶段分割准确率mIoU达0.4881,问答任务最高准确率64.70%。
  • 适合智能交通、自动驾驶系统研发者用于事故复盘与安全优化。

行人-车辆事故仍是城市交通安全的重大挑战,全球交通事故死亡人数中行人占比超20%。现有视频系统虽能检测事故发生,却难以揭示行人行为在不同认知阶段的演化过程。尽管视觉语言模型(VLMs)在视频理解方面展现出潜力,但通常独立处理视频,缺乏显式的时间结构与多视角融合。本文提出多视角阶段感知行人-车辆事故推理框架(MP-PVIR),通过四个阶段实现多视角视频到结构化诊断报告的转化:(1) 事件触发的多视角视频采集,(2) 行人行为阶段分割,(3) 阶段特定的多视角推理,(4) 分层合成与诊断推理。该框架基于行为理论,自动将事故划分为认知阶段,在每个阶段内进行同步多视角分析,并合成因果链与针对性预防策略。其中两个专用VLM支撑流程:TG-VLM用于行为阶段分割(mIoU=0.4881),PhaVR-VLM用于阶段感知多视角分析,生成描述得分33.063,问答任务最高准确率达64.70%。最后由大语言模型生成包含场景理解、行为解释、因果推理与预防建议的完整报告。在Woven Traffic Safety数据集上的评估表明,MP-PVIR能有效将多视角视频转化为可行动洞察,推动车-基础设施协同系统的智能化交通安全分析。

原文摘要 · Abstract (English)

Pedestrian-vehicle incidents remain a critical urban safety challenge, with pedestrians accounting for over 20% of global traffic fatalities. Although existing video-based systems can detect when incidents occur, they provide little insight into how these events unfold across the distinct cognitive phases of pedestrian behavior. Recent vision-language models (VLMs) have shown strong potential for video understanding, but they remain limited in that they typically process videos in isolation, without explicit temporal structuring or multi-view integration. This paper introduces Multi-view Phase-aware Pedestrian-Vehicle Incident Reasoning (MP-PVIR), a unified framework that systematically processes multi-view video streams into structured diagnostic reports through four stages: (1) event-triggered multi-view video acquisition, (2) pedestrian behavior phase segmentation, (3) phase-specific multi-view reasoning, and (4) hierarchical synthesis and diagnostic reasoning. The framework operationalizes behavioral theory by automatically segmenting incidents into cognitive phases, performing synchronized multi-view analysis within each phase, and synthesizing results into causal chains with targeted prevention strategies. Particularly, two specialized VLMs underpin the MP-PVIR pipeline: TG-VLM for behavioral phase segmentation (mIoU = 0.4881) and PhaVR-VLM for phase-aware multi-view analysis, achieving a captioning score of 33.063 and up to 64.70% accuracy on question answering. Finally, a designated large language model is used to generate comprehensive reports detailing scene understanding, behavior interpretation, causal reasoning, and prevention recommendations. Evaluation on the Woven Traffic Safety dataset shows that MP-PVIR effectively translates multi-view video data into actionable insights, advancing AI-driven traffic safety analytics for vehicle-infrastructure cooperative systems.

事故分析多视角视觉语言模型智能交通

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。