arXiv:2507.20146cs.CV2025-07被引 3

不依赖显式对齐,用新方法提升可见光红外目标检测鲁棒性

AlignFreeNet: Is Cross-Modal Pre-Alignment Necessary? An End-to-End Alignment-Free Lightweight Network for Visible-Infrared Object Detection

  • 摒弃传统对齐模块,采用无对齐融合机制
  • 在三个数据集上实现最优性能,尤其在严重混合错位下表现突出
  • 适合复杂环境下的多模态目标检测任务

可见光-红外目标检测(VI-OD)中常存在空间偏移、分辨率差异和语义缺失等跨模态错位问题。现有方法通常引入显式像素或特征级对齐模块以缓解此问题,但像素级对齐难以应对严重或混合错位,特征级对齐在该条件下又会引入噪声,制约检测性能。本文提出一种新型无对齐网络AlignFreeNet,完全摒弃显式对齐,采用无对齐融合范式。其核心包含两个模块:变化引导的跨模态补偿(VCC)与频率引导的跨模态融合(FCF)。VCC通过自适应反馈跨模态差异信息,增强双模态表征且避免对齐噪声;FCF通过频域门控抑制无关冗余,有效降低融合过程中的噪声。二者协同利用低频与高频线索,在融合表示中保留前景轮廓,缓解严重混合错位导致的跨模态混叠。在DVTOD、M3FD和DroneVehicle上的大量实验表明,AlignFreeNet在严重混合错位条件下达到当前最优性能,验证了其鲁棒性与泛化能力。

原文摘要 · Abstract (English)

Cross-modal misalignments, such as spatial offsets, resolution discrepancies, and semantic deficiencies, frequently occur in visible-infrared object detection (VI-OD). To mitigate this, existing methods are typically adapted into an alignment-based fusion paradigm, in which an explicit pixel- or feature-level alignment module is inserted before cross-modal fusion. However, pixel-level alignment struggles to cope with severe or mixed misalignments, whereas feature-level alignment often introduces undesirable noise into fused representations under such conditions, ultimately limiting detection performance. In this paper, we propose a novel alignment-free network (AlignFreeNet) for VI-OD. Differing from prior methods, AlignFreeNet abandons any explicit alignment and instead adopts an alignment-free fusion paradigm. Specifically, AlignFreeNet comprises two core modules: variation-guided cross-modal compensation (VCC) and frequency-guided cross-modal fusion (FCF). VCC adaptively feeds the compensated information derived from cross-modal discrepancies back into each modality, enhancing visible and infrared representations without the noise caused by explicit alignment. FCF achieves robust cross-modal fusion by suppressing task-irrelevant redundancy via frequency-domain gating, effectively mitigating noise introduced in the process. Moreover, VCC and FCF jointly exploit low- and high-frequency cues to preserve foreground contours in fused representations, effectively mitigating cross-modal blending caused by severe mixed misalignments. Extensive evaluations on DVTOD, M3FD, and DroneVehicle demonstrate that our AlignFreeNet achieves state-of-the-art performance under severe mixed misalignment conditions, highlighting its robustness and generalization.

多模态检测无对齐可见光红外轻量网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。