arXiv:2508.21135cs.CVcs.AI2025-08被引 1

融合可见光、热成像和深度信息,提升遮挡目标检测性能

HiddenObject: Modality-Agnostic Fusion for Multimodal Hidden Object Detection

  • 用Mamba架构融合三模态数据,捕捉互补特征
  • 在多个数据集上达到顶尖或领先表现
  • 适合复杂光照与遮挡场景下的目标检测应用

隐藏或部分遮挡物体的检测仍是多模态环境中的核心挑战,遮挡、伪装和光照变化严重影响检测效果。传统基于RGB的方法在恶劣条件下表现不佳,亟需更鲁棒的模态无关方法。本文提出HiddenObject,一种基于Mamba的融合框架,整合RGB、热成像和深度数据。该方法提取各模态特异性特征,并在统一表示中融合,有效增强对被遮挡或伪装目标的检测能力。我们在多个基准数据集上验证了该方法,结果表明其性能优于或媲美现有方法,凸显了所提融合设计的有效性,并揭示了当前单模态及简单融合策略的关键局限。研究进一步表明,基于Mamba的融合架构能显著推动复杂或视觉退化条件下的多模态目标检测发展。

原文摘要 · Abstract (English)

Detecting hidden or partially concealed objects remains a fundamental challenge in multimodal environments, where factors like occlusion, camouflage, and lighting variations significantly hinder performance. Traditional RGB-based detection methods often fail under such adverse conditions, motivating the need for more robust, modality-agnostic approaches. In this work, we present HiddenObject, a fusion framework that integrates RGB, thermal, and depth data using a Mamba-based fusion mechanism. Our method captures complementary signals across modalities, enabling enhanced detection of obscured or camouflaged targets. Specifically, the proposed approach identifies modality-specific features and fuses them in a unified representation that generalizes well across challenging scenarios. We validate HiddenObject across multiple benchmark datasets, demonstrating state-of-the-art or competitive performance compared to existing methods. These results highlight the efficacy of our fusion design and expose key limitations in current unimodal and naïve fusion strategies. More broadly, our findings suggest that Mamba-based fusion architectures can significantly advance the field of multimodal object detection, especially under visually degraded or complex conditions.

多模态融合目标检测Mamba遮挡检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。