arXiv:2411.18288cs.CV2024-11被引 8

首个公平基准+优化技巧包,让多光谱目标检测更稳定高效

Optimizing Multispectral Object Detection: A Bag of Tricks and Comprehensive Benchmarks

  • 构建首个可复现的多光谱检测基准,统一训练配置
  • 验证多种优化技巧在跨数据集上的增益效果,提升检测性能
  • 提供即插即用方案,轻松将单模态强模型转为双模态模型

多光谱目标检测融合可见光(RGB)与热红外(TIR)信息,面临光谱差异、空间错位和环境依赖等挑战,严重影响模型泛化能力。现有研究难以区分性能提升是来自新方法还是优化技巧。尽管单模态检测模型性能强劲,但缺乏针对性的多光谱适配训练策略。为此,本文提出首个公平、可复现的基准,系统分类现有方法,评估超参数敏感性,标准化核心配置。在多个代表性多光谱数据集上,使用不同主干网络和检测框架进行综合评测。同时,提出一种高效易部署的多光谱检测框架,能无缝将高性能单模态模型转化为双模态模型,并集成先进训练技巧。

原文摘要 · Abstract (English)

Multispectral object detection, utilizing RGB and TIR (thermal infrared) modalities, is widely recognized as a challenging task. It requires not only the effective extraction of features from both modalities and robust fusion strategies, but also the ability to address issues such as spectral discrepancies, spatial misalignment, and environmental dependencies between RGB and TIR images. These challenges significantly hinder the generalization of multispectral detection systems across diverse scenarios. Although numerous studies have attempted to overcome these limitations, it remains difficult to clearly distinguish the performance gains of multispectral detection systems from the impact of these "optimization techniques". Worse still, despite the rapid emergence of high-performing single-modality detection models, there is still a lack of specialized training techniques that can effectively adapt these models for multispectral detection tasks. The absence of a standardized benchmark with fair and consistent experimental setups also poses a significant barrier to evaluating the effectiveness of new approaches. To this end, we propose the first fair and reproducible benchmark specifically designed to evaluate the training "techniques", which systematically classifies existing multispectral object detection methods, investigates their sensitivity to hyper-parameters, and standardizes the core configurations. A comprehensive evaluation is conducted across multiple representative multispectral object detection datasets, utilizing various backbone networks and detection frameworks. Additionally, we introduce an efficient and easily deployable multispectral object detection framework that can seamlessly optimize high-performing single-modality models into dual-modality models, integrating our advanced training techniques.

多光谱检测目标检测训练技巧基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。