提出深度协同网络,提升隐蔽物体检测精度。
VCP-DCN: Beyond Visual Concealed Property via Depth Collaborative Network for Camouflaged Object Detection

- 分阶段融合RGB与深度特征,增强多模态表征
- 在三个权威数据集上显著优于现有方法
- 适合研究隐蔽物体检测与多模态融合的学者
隐蔽物体检测(COD)旨在复杂环境中识别并分割外观与背景相似的隐蔽物体。现有方法虽引入深度图以学习互补的RGB-D特征,但忽视了深度域中隐蔽物体的模态特异性。为此,本文提出深度协同网络VCP-DCN,通过挖掘超越视觉隐藏原型的多模态可区分特征,实现更精准的检测。VCP-DCN分三阶段推进:首先,采用可分离原型嵌入(SPE)模块,通过原型对比学习获取一致性和特异性原型令牌;其次,在交互阶段设计多模态双注意力(MDA)模块,利用局部响应图增强跨模态特征表示;最后,在融合阶段设计深度自适应注入(DAI)模块,基于模态特异性与一致性原型令牌间的相似性距离,动态衡量RGB与深度特征贡献。大量实验验证了VCP-DCN在三个权威数据集上的有效性。
原文摘要 · Abstract (English)
Camouflaged Object Detection (COD) aims to identify and segment camouflaged objects in complex environments, which are often concealed because their color and texture are similar to the background. Several existing COD methods introduce depth maps to boost detection performance via learning complementary RGB-D features, ignoring modality-specific characteristics of concealed objects in the depth domain. To address this issue, we propose a depth collaborative network, called VCP-DCN, to mine distinguishable multi-modality features beyond visual concealed prototype in depth domain. Specifically, VCP-DCN progressively performs multi-modality alignment, interaction, and fusion for the COD task. In the \textbf{alignment} stage, we propose a Separable Prototype Embedding (SPE) module to learn modality-consistency and modality-specific RGB/depth prototype tokens through prototype contrastive learning. Furthermore, we develop a Multi-modality Dual Attention (MDA) module to enhance the cross-modal feature representation through local response maps between modality-consistency RGB/depth prototype tokens and visual tokens on the \textbf{interaction} stage. Finally, we design a Depth Adaptive Injection (DAI) module to adaptively measure contribution of RGB/depth features with a decision-making mechanism, which calculates similarity distance between RGB/depth modality-specific prototype tokens and modality-consistency ones on the \textbf{fusion} stage. Extensive experiments demonstrate the effectiveness of our VCP-DCN on three authoritative datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。