arXiv:2510.12190cs.CV2025-10

用分层推理生成行车记录仪事故报告,提升准确性与可读性。

Hierarchical Reasoning with Vision-Language Models for Incident Reports from Dashcam Videos

  • 分帧描述+事件定位+视觉语言模型细粒度推理,层层递进分析视频
  • 在2COOOL挑战赛中排名第二,CIDEr-D得分最优
  • 适合关注自动驾驶安全分析与多模态叙事生成的研究者

端到端自动驾驶的进步依赖于大规模驾驶数据集的训练,但模型在分布外(OOD)场景仍表现不佳。COOOL基准旨在推动对封闭分类体系之外的危险理解,2COOOL挑战赛进一步要求生成人类可读的事故报告。本文提出一种分层推理框架,用于从行车记录仪视频生成事故报告,融合帧级描述、事故帧检测以及视觉语言模型内的细粒度推理。通过模型集成和盲式A/B评分选择协议,进一步提升事实准确性和可读性。在官方2COOOL公开排行榜上,本方法位列29支队伍中的第2名,取得最佳CIDEr-D分数,生成了准确且连贯的事故叙述。结果表明,基于视觉语言模型的分层推理是事故分析及关键交通事件理解的有前景方向。代码已开源:https://github.com/riron1206/kaggle-2COOOL-2nd-Place-Solution。

原文摘要 · Abstract (English)

Recent advances in end-to-end (E2E) autonomous driving have been enabled by training on diverse large-scale driving datasets, yet autonomous driving models still struggle in out-of-distribution (OOD) scenarios. The COOOL benchmark targets this gap by encouraging hazard understanding beyond closed taxonomies, and the 2COOOL challenge extends it to generating human-interpretable incident reports. We present a hierarchical reasoning framework for incident report generation from dashcam videos that integrates frame-level captioning, incident frame detection, and fine-grained reasoning within vision-language models (VLMs). We further improve factual accuracy and readability through model ensembling and a Blind A/B Scoring selection protocol. On the official 2COOOL open leaderboard, our method ranks 2nd among 29 teams and achieves the best CIDEr-D score, producing accurate and coherent incident narratives. These results indicate that hierarchical reasoning with VLMs is a promising direction for accident analysis and for broader understanding of safety-critical traffic events. The implementation and code are available at https://github.com/riron1206/kaggle-2COOOL-2nd-Place-Solution.

事故报告视觉语言模型分层推理自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。