arXiv:2508.09533cs.CVcs.AI2025-08被引 8

提升红外可见光图像中微小目标检测精度,解决错位与多尺度难题。

COXNet: Cross-Layer Fusion with Adaptive Alignment and Scale Integration for RGBT Tiny Object Detection

  • 跨层融合可见光高层特征与红外低层特征,增强语义与空间精度
  • 动态对齐模块纠正模态间空间错位,保留多尺度特征
  • 基于几何形状相似性优化标签分配,适合复杂场景目标检测

在计算机视觉中,从多模态红绿蓝热成像(RGBT)图像中检测微小目标是一项关键挑战,尤其在监控、搜救和自主导航领域。无人机场景因空间错位、低光照、遮挡和背景杂乱加剧了这一挑战。现有方法难以有效利用可见光与热成像之间的互补信息。本文提出COXNet,一种新型的RGBT微小目标检测框架,通过三项核心创新:i)跨层融合模块,融合高层可见光特征与低层热成像特征,提升语义与空间精度;ii)动态对齐与尺度精炼模块,校正跨模态空间错位并保持多尺度特征;iii)基于地理形状相似性度量的优化标签分配策略,改善定位性能。COXNet在RGBTDronePerson数据集上相比当前最优方法,mAP₅₀提升3.32%,验证了其在复杂环境下的鲁棒性。

原文摘要 · Abstract (English)

Detecting tiny objects in multimodal Red-Green-Blue-Thermal (RGBT) imagery is a critical challenge in computer vision, particularly in surveillance, search and rescue, and autonomous navigation. Drone-based scenarios exacerbate these challenges due to spatial misalignment, low-light conditions, occlusion, and cluttered backgrounds. Current methods struggle to leverage the complementary information between visible and thermal modalities effectively. We propose COXNet, a novel framework for RGBT tiny object detection, addressing these issues through three core innovations: i) the Cross-Layer Fusion Module, fusing high-level visible and low-level thermal features for enhanced semantic and spatial accuracy; ii) the Dynamic Alignment and Scale Refinement module, correcting cross-modal spatial misalignments and preserving multi-scale features; and iii) an optimized label assignment strategy using the GeoShape Similarity Measure for better localization. COXNet achieves a 3.32\% mAP$_{50}$ improvement on the RGBTDronePerson dataset over state-of-the-art methods, demonstrating its effectiveness for robust detection in complex environments.

RGBT检测微小目标跨模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。