通过多阶段特征融合与局部注意力机制,提升隐蔽目标检测精度。
Referring Camouflaged Object Detection With Multi-Context Overlapped Windows Cross-Attention
- 多阶段特征融合:结合参考图像与隐蔽目标的深层特征
- 局部注意力机制:在重叠窗口内聚焦细节匹配,提升定位精度
- 适合目标检测与视觉理解研究者,尤其关注隐蔽物体识别
指代式隐蔽目标检测(Ref-COD)旨在结合图像与文本描述等参考信息,定位被隐藏的物体。先前研究将具有显著性目标的参考图像转换为一维提示,取得了显著成果。本文探索通过融合丰富的显著性图像特征与隐蔽目标特征来进一步提升性能。为此,提出RFMNet模型,利用参考显著图像在多个编码阶段的特征,并在对应编码阶段与隐蔽特征进行交互融合。由于显著目标图像中包含大量与物体相关的细节信息,在局部区域内进行特征融合更有利于检测隐蔽目标。因此,提出重叠窗口交叉注意力机制,使模型能基于参考特征更专注地关注局部信息匹配。此外,设计了指代特征聚合(RFA)模块,逐步解码并分割隐蔽目标。在Ref-COD基准数据集上的大量实验表明,本方法达到当前最优性能。
原文摘要 · Abstract (English)
Referring camouflaged object detection (Ref-COD) aims to identify hidden objects by incorporating reference information such as images and text descriptions. Previous research has transformed reference images with salient objects into one-dimensional prompts, yielding significant results. We explore ways to enhance performance through multi-context fusion of rich salient image features and camouflaged object features. Therefore, we propose RFMNet, which utilizes features from multiple encoding stages of the reference salient images and performs interactive fusion with the camouflage features at the corresponding encoding stages. Given that the features in salient object images contain abundant object-related detail information, performing feature fusion within local areas is more beneficial for detecting camouflaged objects. Therefore, we propose an Overlapped Windows Cross-attention mechanism to enable the model to focus more attention on the local information matching based on reference features. Besides, we propose the Referring Feature Aggregation (RFA) module to decode and segment the camouflaged objects progressively. Extensive experiments on the Ref-COD benchmark demonstrate that our method achieves state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。