arXiv:2603.24355cs.CVcs.AI2026-03

用文字提示引导网络聚焦隐蔽目标,提升边界精度。

Language-Guided Structure-Aware Network for Camouflaged Object Detection

  • 引入CLIP生成文本引导掩码,指导多尺度特征聚焦目标区域
  • 通过频域增强提取边缘特征,在多个数据集上达到领先性能
  • 适合需要精准分割隐蔽物体的场景,如生物探测、军事识别

隐蔽目标检测(COD)旨在分割在颜色、纹理和结构上与背景高度融合的物体,是计算机视觉中的难题。现有方法虽采用多尺度融合与注意力机制缓解该问题,但普遍缺乏文本语义先验引导,限制了模型在复杂场景中对隐蔽区域的关注能力。本文提出语言引导结构感知网络(LGSAN),基于视觉主干PVT-v2,利用CLIP从文本提示与图像中生成掩码,引导PVT-v2提取的多尺度特征聚焦潜在目标区域。在此基础上,设计傅里叶边缘增强模块(FEEM),将多尺度特征与频域高频信息结合,提取边缘增强特征;提出结构感知注意力模块(SAAM),有效增强模型对物体结构与边界的感知能力;最后引入粗粒度引导局部精修模块(CGLRM),提升隐蔽目标区域的细粒度重建与边界完整性。大量实验表明,本方法在多个COD数据集上持续表现优异,验证了其有效性与鲁棒性。

原文摘要 · Abstract (English)

Camouflaged Object Detection (COD) aims to segment objects that are highly integrated with the background in terms of color, texture, and structure, making it a highly challenging task in computer vision. Although existing methods introduce multi-scale fusion and attention mechanisms to alleviate the above issues, they generally lack the guidance of textual semantic priors, which limits the model's ability to focus on camouflaged regions in complex scenes. To address this issue, this paper proposes a Language-Guided Structure-Aware Network (LGSAN). Specifically, based on the visual backbone PVT-v2, we introduce CLIP to generate masks from text prompts and RGB images, thereby guiding the multi-scale features extracted by PVT-v2 to focus on potential target regions. On this foundation, we further design a Fourier Edge Enhancement Module (FEEM), which integrates multi-scale features with high-frequency information in the frequency domain to extract edge enhancement features. Furthermore, we propose a Structure-Aware Attention Module (SAAM) to effectively enhance the model's perception of object structures and boundaries. Finally, we introduce a Coarse-Guided Local Refinement Module (CGLRM) to enhance fine-grained reconstruction and boundary integrity of camouflaged object regions. Extensive experiments demonstrate that our method consistently achieves highly competitive performance across multiple COD datasets, validating its effectiveness and robustness.

隐蔽目标检测视觉语言模型边缘增强结构感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。