arXiv:2607.03591cs.CVcs.AI2026-07

用多模态大模型分析行车视角事故视频,估算各方责任比例。

Responsibility Distribution Estimation in Ego-View Accident Videos with Multimodal Large Language Models

  • 构建基于大模型的标注流程,融合图像与文本输入进行责任分配。
  • 在多模态输入下,模型能有效预测各涉事方的责任占比,性能显著。
  • 适合交通责任判定、自动驾驶安全评估等需要主观推理的场景。

现有交通事故理解研究多依赖基础设施摄像头、卫星影像或结构化事故记录,但这些数据部署维护成本高,且无法客观反映驾驶员事发前的实际视觉感知。相比之下,行车视角视频直接呈现驾驶员的视觉视角,更适用于判断事故可避免性及责任归属。本文提出行车视角事故视频中的责任分布估计新任务,即模型需预测各涉事方的责任占比。我们构建了基于大模型的辅助标注流程,并在多种输入设置(原始帧、分割增强输入、文本描述)下微调多模态大语言模型。实验建立了首个基准,表明多模态大模型能有效完成这一具有约束条件的细微推理任务。结果表明,以驾驶员为中心的事故视频为超越传统分类与解释任务的社会化、法律化多模态推理提供了良好基础。

原文摘要 · Abstract (English)

Recent studies on multimodal traffic accident understanding have mainly relied on infrastructure-camera footage, satellite imagery, or structured crash records. However, such data sources are costly to deploy and maintain at large scale, and they cannot objectively capture what the driver was actually able to observe before the accident. In contrast, ego-view accident videos directly represent the driver's visual perspective, making them suitable for reasoning about avoidability and driver responsibility. In this paper, we introduce responsibility distribution estimation for ego-view traffic accident videos, a new task in which a model predicts the percentage of responsibility assigned to each involved agent. We construct an LLM-assisted responsibility annotation pipeline and fine-tune multimodal large language models under multiple input settings, including raw frames, segmentation-enhanced input, and textual descriptions. Experimental results establish a strong initial benchmark, demonstrating that multimodal LLMs can effectively perform this nuanced, constraint-based reasoning task. Our findings suggest that ego-centric accident videos provide a promising foundation for socially and legally meaningful multimodal reasoning beyond conventional accident classification and explanation tasks.

事故分析多模态大模型责任判定

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。