用视频和法律规则自动判断交通事故责任,透明可解释。
Interpretable Traffic Responsibility from Dashcam Video via Legal Multi Agent Reasoning
- 分两阶段:先生成视频文字描述,再通过多智能体推理责任归属。
- 在两个数据集上超越现有模型,准确率显著提升。
- 适合法律AI、交通管理及自动驾驶事故判定研究者。
行车记录仪的普及使交通事故视频证据日益丰富,但将“视频中发生了什么”转化为“依据哪条法规应由谁担责”,仍严重依赖人工专家。现有基于单视角视频的研究集中于感知与语义理解,而基于大模型的法律方法多依赖文本案例描述,极少结合视频证据,导致二者之间存在明显断层。本文首次提出C-TRAIL,一个面向中国交通法规体系的多模态法律数据集,明确对齐行车记录视频、文本描述与责任归类及其对应的中文交通法规条文。在此基础上,构建两阶段框架:(1) 事故理解模块生成视频的文本描述;(2) 法律多智能体框架输出责任模式、法规集合及完整判决报告。在C-TRAIL与MM-AU数据集上的实验表明,该方法优于通用及法律大模型,以及现有基于智能体的方法,同时提供透明、可解释的法律推理过程。
原文摘要 · Abstract (English)
The widespread adoption of dashcams has made video evidence in traffic accidents increasingly abundant, yet transforming "what happened in the video" into "who is responsible under which legal provisions" still relies heavily on human experts. Existing ego-view traffic accident studies mainly focus on perception and semantic understanding, while LLM-based legal methods are mostly built on textual case descriptions and rarely incorporate video evidence, leaving a clear gap between the two. We first propose C-TRAIL, a multimodal legal dataset that, under the Chinese traffic regulation system, explicitly aligns dashcam videos and textual descriptions with a closed set of responsibility modes and their corresponding Chinese traffic statutes. On this basis, we introduce a two-stage framework: (1) a traffic accident understanding module that generates textual video descriptions; and (2) a legal multi-agent framework that outputs responsibility modes, statute sets, and complete judgment reports. Experimental results on C-TRAIL and MM-AU show that our method outperforms general and legal LLMs, as well as existing agent-based approaches, while providing a transparent and interpretable legal reasoning process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。