arXiv:2506.07779cs.CV2025-06

构建校园场景双模态数据集,评估融合模型在目标检测中的实际表现。

Design and Evaluation of Deep Learning-Based Dual-Spectrum Image Fusion Methods

  • 基于校园环境构建1369对可见光与红外图像数据集。
  • 发现通用指标表现好的模型在目标检测中未必有效。
  • 提出结合检测任务的公平评估框架,适合下游应用研究者。

可见光图像富含纹理细节,红外图像突出显著目标。融合这两种互补模态可增强复杂条件下的场景理解,尤其适用于高级视觉任务。近年来深度学习融合方法受关注,但现有评估多依赖通用指标,缺乏标准化基准和下游任务性能验证。同时,高质量双模态数据集稀缺且算法比较不公,制约发展。为此,我们构建了校园环境下的高质双模态数据集,包含1369对配准良好的可见光-红外图像,覆盖白天、夜晚、烟雾遮挡和地下通道四种典型场景。提出一个综合评估框架,融合融合速度、通用指标及使用lang-segment-anything模型的目标检测性能,确保下游评估公平性。大量实验在该框架下对比多个先进融合算法。结果表明,针对下游任务优化的模型在低光照和遮挡场景中检测性能更优。值得注意的是,部分通用指标表现优异的算法在下游任务中表现不佳,揭示当前评估方式局限性,并验证本框架必要性。主要贡献:(1) 面向校园场景的多样化挑战性双模态数据集;(2) 任务感知的综合性评估框架;(3) 在多数据集上对主流融合方法的全面分析,为未来研究提供洞见。

原文摘要 · Abstract (English)

Visible images offer rich texture details, while infrared images emphasize salient targets. Fusing these complementary modalities enhances scene understanding, particularly for advanced vision tasks under challenging conditions. Recently, deep learning-based fusion methods have gained attention, but current evaluations primarily rely on general-purpose metrics without standardized benchmarks or downstream task performance. Additionally, the lack of well-developed dual-spectrum datasets and fair algorithm comparisons hinders progress. To address these gaps, we construct a high-quality dual-spectrum dataset captured in campus environments, comprising 1,369 well-aligned visible-infrared image pairs across four representative scenarios: daytime, nighttime, smoke occlusion, and underpasses. We also propose a comprehensive and fair evaluation framework that integrates fusion speed, general metrics, and object detection performance using the lang-segment-anything model to ensure fairness in downstream evaluation. Extensive experiments benchmark several state-of-the-art fusion algorithms under this framework. Results demonstrate that fusion models optimized for downstream tasks achieve superior performance in target detection, especially in low-light and occluded scenes. Notably, some algorithms that perform well on general metrics do not translate to strong downstream performance, highlighting limitations of current evaluation practices and validating the necessity of our proposed framework. The main contributions of this work are: (1)a campus-oriented dual-spectrum dataset with diverse and challenging scenes; (2) a task-aware, comprehensive evaluation framework; and (3) thorough comparative analysis of leading fusion methods across multiple datasets, offering insights for future development.

图像融合多模态目标检测数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。