arXiv:2503.17050cs.CV2025-03ICCV被引 2

通过记忆参考帧提升视频中伪装目标检测精度

Scoring, Remember, and Reference: Catching Camouflaged Objects in Videos

  • 借鉴人类记忆机制,用评分选参考帧并动态更新
  • 在基准数据集上性能领先10%,参数仅54M
  • 适合需要高效精准视频目标检测的场景

视频伪装目标检测(VCOD)旨在分割外观与环境高度相似的目标,是极具挑战性的新兴任务。现有视觉模型因目标与背景难以区分、视频动态信息利用不足而表现不佳。为此,我们提出一种受人类记忆-识别机制启发的端到端VCOD框架,通过引入历史视频帧作为记忆参考,实现对伪装序列的有效处理。设计双功能解码器,同时生成预测掩码与置信度分数,依据分数选择参考帧,并引入辅助监督增强特征提取。此外,提出新型参考引导的多层级非对称注意力机制,有效融合长时参考信息与短时运动线索,实现全面特征建模。结合上述模块,构建得分、记忆与参考(SRR)框架,高效提取目标信息并利用记忆指导后续处理。该模型在基准数据集上性能提升10%,仅需单次视频遍历,参数量为54M,代码将公开。

原文摘要 · Abstract (English)

Video Camouflaged Object Detection (VCOD) aims to segment objects whose appearances closely resemble their surroundings, posing a challenging and emerging task. Existing vision models often struggle in such scenarios due to the indistinguishable appearance of camouflaged objects and the insufficient exploitation of dynamic information in videos. To address these challenges, we propose an end-to-end VCOD framework inspired by human memory-recognition, which leverages historical video information by integrating memory reference frames for camouflaged sequence processing. Specifically, we design a dual-purpose decoder that simultaneously generates predicted masks and scores, enabling reference frame selection based on scores while introducing auxiliary supervision to enhance feature extraction.Furthermore, this study introduces a novel reference-guided multilevel asymmetric attention mechanism, effectively integrating long-term reference information with short-term motion cues for comprehensive feature extraction. By combining these modules, we develop the Scoring, Remember, and Reference (SRR) framework, which efficiently extracts information to locate targets and employs memory guidance to improve subsequent processing. With its optimized module design and effective utilization of video data, our model achieves significant performance improvements, surpassing existing approaches by 10% on benchmark datasets while requiring fewer parameters (54M) and only a single pass through the video. The code will be made publicly available.

视频检测伪装目标记忆机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。