通过历史记忆增强,提升隐蔽目标检测在复杂场景下的感知能力。
Retrospective Memory for Camouflaged Object Detection
- 引入动态记忆机制,融合历史知识优化当前推理
- 在多个数据集上超越现有最先进方法,显著提升检测性能
- 适合研究视觉记忆与目标检测结合的学者参考
隐蔽目标检测(COD)主要关注从复杂场景中学习细微但具有区分性的表征。现有方法多采用基于静态视觉表征建模的参数化前馈架构,缺乏获取历史上下文的显式机制,限制了其在挑战性伪装场景中的适应性和有效性。本文提出一种回忆增强型COD架构RetroMem,通过将相关历史知识融入推理过程,动态调节伪装模式感知与判断。具体地,RetroMem采用两阶段训练范式:学习阶段设计密集多尺度适配器(DMA),以极少可训练参数增强预训练编码器捕捉丰富多尺度视觉信息的能力,提供基础推理;召回阶段提出动态记忆机制(DMM)与推理模式重建(IPR),充分挖掘已学知识与当前样本上下文间的潜在关系,重构伪装模式推理,显著提升模型对伪装场景的理解能力。大量实验表明,RetroMem在多个常用数据集上显著优于现有最先进方法。
原文摘要 · Abstract (English)
Camouflaged object detection (COD) primarily focuses on learning subtle yet discriminative representations from complex scenes. Existing methods predominantly follow the parametric feedforward architecture based on static visual representation modeling. However, they lack explicit mechanisms for acquiring historical context, limiting their adaptation and effectiveness in handling challenging camouflage scenes. In this paper, we propose a recall-augmented COD architecture, namely RetroMem, which dynamically modulates camouflage pattern perception and inference by integrating relevant historical knowledge into the process. Specifically, RetroMem employs a two-stage training paradigm consisting of a learning stage and a recall stage to construct, update, and utilize memory representations effectively. During the learning stage, we design a dense multi-scale adapter (DMA) to improve the pretrained encoder's capability to capture rich multi-scale visual information with very few trainable parameters, thereby providing foundational inferences. In the recall stage, we propose a dynamic memory mechanism (DMM) and an inference pattern reconstruction (IPR). These components fully leverage the latent relationships between learned knowledge and current sample context to reconstruct the inference of camouflage patterns, thereby significantly improving the model's understanding of camouflage scenes. Extensive experiments on several widely used datasets demonstrate that our RetroMem significantly outperforms existing state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。