arXiv:2601.18008cs.CV2026-01中稿 · ICRA被引 3

提出时空融合网络,提升多光谱行人检测在错位与遮挡下的性能。

Strip-Fusion: Spatiotemporal Fusion for Multispectral Pedestrian Detection

  • 引入时序自适应卷积,动态融合时空特征以捕捉运动信息。
  • 在KAIST和CVC-14上优于现有方法,尤其在严重遮挡与图像错位时提升显著。
  • 适合需要鲁棒视觉感知的机器人、自动驾驶场景使用。

行人检测是机器人感知的关键任务。多光谱模态(可见光与热成像)可通过互补视觉信息提升检测性能。现有方法主要关注空间融合,常忽略时间信息;且基准数据集中可见光与热成像对未必完全对齐。行人检测还面临光照变化、遮挡等挑战。本文提出Strip-Fusion,一种对输入图像错位、光照变化及严重遮挡具有鲁棒性的时空融合网络。该框架采用时序自适应卷积,动态加权时空特征,更好捕捉行人运动与上下文。设计了一种新型Kullback-Leibler散度损失,缓解可见光与热成像之间的模态不平衡,训练时引导特征对齐至更有效模态。此外,提出新后处理算法降低误报率。大量实验表明,该方法在KAIST与CVC-14基准上表现优异,尤其在重遮挡与错位条件下较先前最先进方法有显著提升。

原文摘要 · Abstract (English)

Pedestrian detection is a critical task in robot perception. Multispectral modalities (visible light and thermal) can boost pedestrian detection performance by providing complementary visual information. Several gaps remain with multispectral pedestrian detection methods. First, existing approaches primarily focus on spatial fusion and often neglect temporal information. Second, RGB and thermal image pairs in multispectral benchmarks may not always be perfectly aligned. Pedestrians are also challenging to detect due to varying lighting conditions, occlusion, etc. This work proposes Strip-Fusion, a spatial-temporal fusion network that is robust to misalignment in input images, as well as varying lighting conditions and heavy occlusions. The Strip-Fusion pipeline integrates temporally adaptive convolutions to dynamically weigh spatial-temporal features, enabling our model to better capture pedestrian motion and context over time. A novel Kullback-Leibler divergence loss was designed to mitigate modality imbalance between visible and thermal inputs, guiding feature alignment toward the more informative modality during training. Furthermore, a novel post-processing algorithm was developed to reduce false positives. Extensive experimental results show that our method performs competitively for both the KAIST and the CVC-14 benchmarks. We also observed significant improvements compared to previous state-of-the-art on challenging conditions such as heavy occlusion and misalignment.

行人检测多光谱时空融合鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。