arXiv:2504.03235cs.CVcs.AI2025-04被引 1

用混合模型实现车祸秒级定位,提升应急响应效率

Enhancing Traffic Incident Response through Sub-Second Temporal Localization with HybridMamba

  • 融合视觉变换器与状态空间模型,分层压缩保持时间精度
  • 2分钟视频定位误差仅1.50秒,65.2%预测在1秒内
  • 参数量仅30亿,远低于同类模型,适合实时部署

长时监控视频中的交通事故检测对提升应急响应和基础设施规划至关重要,但因事故短暂且罕见而极具挑战。我们提出 HybridMamba,一种结合视觉变换器与状态空间时间建模的新型架构,实现高精度事故时间定位。通过多层次标记压缩与分层时间处理,在保持计算效率的同时不损失时间分辨率。在爱荷华州交通部的大规模数据集上评估,HybridMamba 在2分钟视频上的平均绝对误差为1.50秒(与基线相比,p<0.01),65.2%的预测结果在真实时刻1秒内。其性能优于近期视频-语言模型(如 TimeChat、VideoLLaMA-2),最大误差降低达3.95秒,同时仅使用30亿参数(对比13–720亿),在多种视频时长(2–40分钟)和环境条件下均表现良好,展示了其在交通监控中精细时间定位的潜力,并指出了未来扩展部署仍面临挑战。

原文摘要 · Abstract (English)

Traffic crash detection in long-form surveillance videos is essential for improving emergency response and infrastructure planning, yet remains difficult due to the brief and infrequent nature of crash events. We present \textbf{HybridMamba}, a novel architecture integrating visual transformers with state-space temporal modeling to achieve high-precision crash time localization. Our approach introduces multi-level token compression and hierarchical temporal processing to maintain computational efficiency without sacrificing temporal resolution. Evaluated on a large-scale dataset from the Iowa Department of Transportation, HybridMamba achieves a mean absolute error of \textbf{1.50 seconds} for 2-minute videos ($p<0.01$ compared to baselines), with \textbf{65.2%} of predictions falling within one second of the ground truth. It outperforms recent video-language models (e.g., TimeChat, VideoLLaMA-2) by up to 3.95 seconds while using significantly fewer parameters (3B vs. 13--72B). Our results demonstrate effective temporal localization across various video durations (2--40 minutes) and diverse environmental conditions, highlighting HybridMamba's potential for fine-grained temporal localization in traffic surveillance while identifying challenges that remain for extended deployment.

事故检测时间定位轻量化模型视频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。