arXiv:2502.18496cs.CV2025-02

融合深度信息与多维视觉特征,提升行车事故早期预测能力

Physical Depth-aware Early Accident Anticipation: A Multi-dimensional Visual Feature Fusion Framework

  • 引入单目深度图增强3D空间感知,弥补2D图像信息不足
  • 在公开数据集上达到当前最优性能,显著提升预测准确率
  • 适合智能驾驶系统开发、交通安全研究者参考

从行车记录仪视频中提前预测事故是提升智能汽车安全性的重要任务,但现有方法多在粗粒度的2D图像空间建模交通参与者(如车辆、行人)间的交互关系,难以准确捕捉其真实位置与互动。为此,本文提出一种物理深度感知的学习框架,利用名为Depth-Anything的大模型生成的单目深度特征,引入更精细的三维空间信息。同时,该框架还融合了交通场景中的视觉交互特征与动态特征,实现对场景的更全面感知。基于这些多维度视觉特征,框架通过分析连续帧间物体之间的交互关系,捕捉事故的早期信号。此外,针对被遮挡的关键交通参与者,框架设计了一种重构邻接矩阵,缓解遮挡对图学习的影响,保持时空连续性。在多个公开数据集上的实验表明,该框架达到当前最优性能,验证了视觉深度特征的有效性及框架的优势。

原文摘要 · Abstract (English)

Early accident anticipation from dashcam videos is a highly desirable yet challenging task for improving the safety of intelligent vehicles. Existing advanced accident anticipation approaches commonly model the interaction among traffic agents (e.g., vehicles, pedestrians, etc.) in the coarse 2D image space, which may not adequately capture their true positions and interactions. To address this limitation, we propose a physical depth-aware learning framework that incorporates the monocular depth features generated by a large model named Depth-Anything to introduce more fine-grained spatial 3D information. Furthermore, the proposed framework also integrates visual interaction features and visual dynamic features from traffic scenes to provide a more comprehensive perception towards the scenes. Based on these multi-dimensional visual features, the framework captures early indicators of accidents through the analysis of interaction relationships between objects in sequential frames. Additionally, the proposed framework introduces a reconstruction adjacency matrix for key traffic participants that are occluded, mitigating the impact of occluded objects on graph learning and maintaining the spatio-temporal continuity. Experimental results on public datasets show that the proposed framework attains state-of-the-art performance, highlighting the effectiveness of incorporating visual depth features and the superiority of the proposed framework.

事故预测深度感知智能驾驶多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。