改进红外小目标检测评估方法,提升模型分析与泛化能力
Rethinking Evaluation of Infrared Small Target Detection
- 融合像素级与目标级指标,构建混合评估框架
- 提出系统性错误分析方法,揭示模型失效模式
- 倡导跨数据集测试,推动模型泛化能力研究
红外小目标检测(IRSTD)作为重要视觉任务,虽因深度学习取得显著进展,但现有评估协议存在关键缺陷:其一,现有方法依赖碎片化的像素级与目标级指标,难以全面反映模型性能;其二,过度关注整体性能得分,掩盖了关键错误分析,不利于识别失败模式与提升实际系统表现;其三,普遍采用数据集特定的训练-测试范式,阻碍对模型在多样化红外场景下鲁棒性与泛化能力的理解。本文通过引入融合像素与目标层级的混合指标、提出系统性错误分析方法,并强调跨数据集评估的重要性,旨在构建更全面、合理的分层分析框架,推动更高效、更鲁棒的IRSTD模型发展。相关开源工具包已发布,以支持标准化基准测试。
原文摘要 · Abstract (English)
As an essential vision task, infrared small target detection (IRSTD) has seen significant advancements through deep learning. However, critical limitations in current evaluation protocols impede further progress. First, existing methods rely on fragmented pixel- and target-level specific metrics, which fails to provide a comprehensive view of model capabilities. Second, an excessive emphasis on overall performance scores obscures crucial error analysis, which is vital for identifying failure modes and improving real-world system performance. Third, the field predominantly adopts dataset-specific training-testing paradigms, hindering the understanding of model robustness and generalization across diverse infrared scenarios. This paper addresses these issues by introducing a hybrid-level metric incorporating pixel- and target-level performance, proposing a systematic error analysis method, and emphasizing the importance of cross-dataset evaluation. These aim to offer a more thorough and rational hierarchical analysis framework, ultimately fostering the development of more effective and robust IRSTD models. An open-source toolkit has be released to facilitate standardized benchmarking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。