arXiv:2604.16240cs.CV2026-04中稿 · ICPR 2026

CollideNet通过分层多尺度建模,提升碰撞时间预测精度。

CollideNet: Hierarchical Multi-scale Video Representation Learning with Disentanglement for Time-To-Collision Forecasting

论文配图:CollideNet: Hierarchical Multi-scale Video Representation Learning with Disentanglement for Time-To-Collision Forecasting
图 1 · 摘自论文原文
  • 分层时空架构,多分辨率融合视频帧信息
  • 在三个数据集上达到最新最佳性能,显著超越此前方法
  • 解耦趋势与季节性成分,适合自动驾驶场景研究

碰撞时间(TTC)预测是预防碰撞的关键任务,需精确的时间预测,并理解视频中时空层面的局部与全局模式。为应对视频的多尺度特性,我们提出一种新型基于时空分层变换器的架构CollideNet,专为高效TTC预测设计。空间流中,CollideNet在多个分辨率下同时聚合每帧信息;时间流中,除多尺度特征编码外,还解耦非平稳性、趋势与季节性成分。该方法在三个常用公开数据集上表现优于先前工作,显著提升性能。我们进行了跨数据集评估以分析泛化能力,并可视化了趋势与季节性解耦效果。代码已开源:https://github.com/DeSinister/CollideNet/

原文摘要 · Abstract (English)

Time-to-Collision (TTC) forecasting is a critical task in collision prevention, requiring precise temporal prediction and comprehending both local and global patterns encapsulated in a video, both spatially and temporally. To address the multi-scale nature of video, we introduce a novel spatiotemporal hierarchical transformer-based architecture called CollideNet, specifically catered for effective TTC forecasting. In the spatial stream, CollideNet aggregates information for each video frame simultaneously at multiple resolutions. In the temporal stream, along with multi-scale feature encoding, CollideNet also disentangles the non-stationarity, trend, and seasonality components. Our method achieves state-of-the-art performance in comparison to prior works on three commonly used public datasets, setting a new state-of-the-art by a considerable margin. We conduct cross-dataset evaluations to analyze the generalization capabilities of our method, and visualize the effects of disentanglement of the trend and seasonality components of the video data. We release our code at https://github.com/DeSinister/CollideNet/.

视频预测多尺度建模自动驾驶时空建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。