通过多尺度特征融合提升行车事故预测准确性和提前量
MsFIN: Multi-scale Feature Interaction Network for Traffic Accident Anticipation
- 设计多尺度模块提取短中长期场景特征,结合Transformer增强交互
- 在DAD和DADA数据集上,预测准确率与提前量均优于现有模型
- 适合自动驾驶安全预警系统开发人员参考
随着行车记录仪普及和计算机视觉发展,从行车视角进行事故预测对主动安全干预至关重要。但两大挑战仍存:难以建模交通参与者间(常被遮挡)的特征级交互,以及捕捉复杂、异步的多时序行为线索。为此提出多尺度特征交互网络(MsFIN),包含三层:多尺度特征聚合、时序特征处理与多尺度特征后融合。多尺度模块在短、中、长时序尺度提取场景表征,结合Transformer实现全面特征交互;时序处理在因果约束下捕获场景与物体特征的演化;后融合阶段跨时序融合场景与物体特征,生成综合风险表征。在DAD和DADA数据集上的实验表明,MsFIN显著优于单尺度特征提取的先进模型,在预测正确率与提前量上均有提升。消融实验证实各模块有效性,凸显多尺度融合与上下文交互建模的关键作用。
原文摘要 · Abstract (English)
With the widespread deployment of dashcams and advancements in computer vision, developing accident prediction models from the dashcam perspective has become critical for proactive safety interventions. However, two key challenges persist: modeling feature-level interactions among traffic participants (often occluded in dashcam views) and capturing complex, asynchronous multi-temporal behavioral cues preceding accidents. To deal with these two challenges, a Multi-scale Feature Interaction Network (MsFIN) is proposed for early-stage accident anticipation from dashcam videos. MsFIN has three layers for multi-scale feature aggregation, temporal feature processing and multi-scale feature post fusion, respectively. For multi-scale feature aggregation, a Multi-scale Module is designed to extract scene representations at short-term, mid-term and long-term temporal scales. Meanwhile, the Transformer architecture is leveraged to facilitate comprehensive feature interactions. Temporal feature processing captures the sequential evolution of scene and object features under causal constraints. In the multi-scale feature post fusion stage, the network fuses scene and object features across multiple temporal scales to generate a comprehensive risk representation. Experiments on DAD and DADA datasets show that MsFIN significantly outperforms state-of-the-art models with single-scale feature extraction in both prediction correctness and earliness. Ablation studies validate the effectiveness of each module in MsFIN, highlighting how the network achieves superior performance through multi-scale feature fusion and contextual interaction modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。