arXiv:2608.07941cs.CV2026-08

让语言指令精准指导视觉细节,提升隐身物体检测精度

LAD-COD: Language-Aligned Dense Perception for Camouflaged Object Detection

论文配图:LAD-COD: Language-Aligned Dense Perception for Camouflaged Object Detection
图 1 · 摘自论文原文
  • 用语言指令引导底层视觉特征,实现语义与细节的双向对齐
  • 在3个数据集上12项指标全领先,最高提升4.6%
  • 适合需要高精度定位复杂伪装目标的研究与应用

隐身物体检测(COD)旨在分割与背景视觉高度相似、难以区分的物体,此类物体在外观、纹理和结构上边界证据弱。现有基于大模型的系统虽能通过指令生成目标嵌入引导分割,但仅作用于解码器,未对密集视觉特征进行显式语义引导。为此提出LAD-COD框架,将自上而下的语言语义与自下而上的层级视觉特征对齐。不直接复用通用图像编码器,而是学习可训练的层次化视觉分支,捕捉敏感于伪装的纹理、边界和上下文信息。通过语言对齐双路视觉融合(LADVF),将目标嵌入扩展至块级特征查询,并控制其残差融合,使语义信息有效引导定位同时保留精细结构。在CAMO、COD10K和NC4K数据集上的实验表明,LAD-COD在全部12项指标中均取得最佳表现。

原文摘要 · Abstract (English)

Camouflaged object detection (COD) aims to segment objects that exhibit high visual similarity to their surroundings, which reduces foreground-background discriminability and weakens boundary evidence across appearance, texture, and structure. Such limitations motivate the use of instruction-conditioned semantics as top-down guidance for identifying which weak visual cues are relevant to the target. Recent segmentation systems built on large multimodal models (LMMs) demonstrate this possibility through instruction-conditioned target embeddings that guide mask decoding. However, in this language-to-mask paradigm, the generated target embedding conditions mainly the mask decoder, leaving the dense visual features that must preserve low-contrast boundaries and fine local structure without explicit guidance. We propose Language-Aligned Dense perception for COD (LAD-COD), a framework that aligns top-down semantic target guidance with bottom-up hierarchical visual features. Instead of fully adapting a large generic image encoder, LAD-COD learns a trainable hierarchical visual branch that captures camouflage-sensitive texture, boundary, and contextual information. To align these features with the target embedding, LAD-COD applies Language-Aligned Dual Visual Fusion (LADVF), which extends the embedding beyond sparse prompting to query patch-level language-aligned features and to gate their residual integration with the hierarchical features. This design allows semantic information to guide localization while preserving the fine structural details needed for camouflage segmentation. Experiments on CAMO, COD10K, and NC4K show that LAD-COD obtains the best reported value in all 12 dataset-metric comparisons.

隐身物体检测语言对齐视觉融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。