arXiv:2507.12755cs.CVcs.LG2025-07被引 9

融合视频与事故报告文本,提升自动驾驶事故预测精度与可解释性。

Domain-Enhanced Dual-Branch Model for Efficient and Interpretable Accident Anticipation

  • 双分支结构分别处理摄像头视频和事故报告文本,实现多模态融合。
  • 在DAD、CCD、A3D数据集上准确率领先,计算开销更低。
  • 结合大模型与提示工程,输出可操作建议和标准化档案,适合安全系统研发者。

构建精准且计算高效的交通事故预测系统对现代自动驾驶技术至关重要,可实现及时干预并减少损失。本文提出一种基于双分支架构的事故预测框架,有效融合行车记录仪视频的视觉信息与事故报告中的结构化文本数据。同时引入特征聚合方法,通过大模型(GPT-4o、Long-CLIP)实现多模态输入的无缝整合,并结合针对性提示工程策略,生成可操作反馈与标准化事故档案。在基准数据集DAD、CCD和A3D上的综合评估表明,该方法在预测准确性、响应速度、计算开销和可解释性方面均表现优越,确立了当前交通事故预测领域的最新性能标杆。

原文摘要 · Abstract (English)

Developing precise and computationally efficient traffic accident anticipation system is crucial for contemporary autonomous driving technologies, enabling timely intervention and loss prevention. In this paper, we propose an accident anticipation framework employing a dual-branch architecture that effectively integrates visual information from dashcam videos with structured textual data derived from accident reports. Furthermore, we introduce a feature aggregation method that facilitates seamless integration of multimodal inputs through large models (GPT-4o, Long-CLIP), complemented by targeted prompt engineering strategies to produce actionable feedback and standardized accident archives. Comprehensive evaluations conducted on benchmark datasets (DAD, CCD, and A3D) validate the superior predictive accuracy, enhanced responsiveness, reduced computational overhead, and improved interpretability of our approach, thus establishing a new benchmark for state-of-the-art performance in traffic accident anticipation.

事故预测多模态可解释性自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。