arXiv:2603.26109cs.CV2026-03中稿 · CVPR

针对伪装物体检测难题,提出动态聚焦方法提升识别能力。

SDDF: Specificity-Driven Dynamic Focusing for Open-Vocabulary Camouflaged Object Detection

论文配图:SDDF: Specificity-Driven Dynamic Focusing for Open-Vocabulary Camouflaged Object Detection
图 1 · 摘自论文原文
  • 用细粒度文本描述增强伪装物体数据,构建新基准OVCOD-D
  • 通过主成分对比融合降低文本噪声,提升语义准确性
  • 设计特定性引导的动态聚焦机制,有效区分物体与背景

开放词汇目标检测(OVOD)旨在利用文本提示检测开放世界中已知和未知物体。得益于大规模视觉-语言预训练模型的发展,OVOD展现出强大的零样本泛化能力。然而,在处理伪装物体时,由于物体与背景视觉特征高度相似,检测器常无法准确区分和定位。为此,我们构建了名为OVCOD-D的新基准,通过精心选择的伪装物体图像并添加细粒度文本描述。由于现有伪装物体数据集规模有限,我们采用在大规模目标检测数据集上预训练的检测器作为基线方法,因其具备更强的零样本泛化能力。在多模态大模型生成的细粒度子描述中,仍存在混淆性及过度修饰的修饰词。为缓解此类干扰,我们设计了子描述主成分对比融合策略,以减少噪声文本成分。此外,针对伪装物体视觉特征与环境高度相似的挑战,提出一种特定性引导的区域弱对齐与动态聚焦方法,旨在增强检测器区分物体与背景的能力。在开放集评估设置下,所提方法在OVCOD-D基准上达到56.4的AP值。

原文摘要 · Abstract (English)

Open-vocabulary object detection (OVOD) aims to detect known and unknown objects in the open world by leveraging text prompts. Benefiting from the emergence of large-scale vision--language pre-trained models, OVOD has demonstrated strong zero-shot generalization capabilities. However, when dealing with camouflaged objects, the detector often fails to distinguish and localize objects because the visual features of the objects and the background are highly similar. To bridge this gap, we construct a benchmark named OVCOD-D by augmenting carefully selected camouflaged object images with fine-grained textual descriptions. Due to the limited scale of available camouflaged object datasets, we adopt detectors pre-trained on large-scale object detection datasets as our baseline methods, as they possess stronger zero-shot generalization ability. In the specificity-aware sub-descriptions generated by multimodal large models, there still exist confusing and overly decorative modifiers. To mitigate such interference, we design a sub-description principal component contrastive fusion strategy that reduces noisy textual components. Furthermore, to address the challenge that the visual features of camouflaged objects are highly similar to those of their surrounding environment, we propose a specificity-guided regional weak alignment and dynamic focusing method, which aims to strengthen the detector's ability to discriminate camouflaged objects from background. Under the open-set evaluation setting, the proposed method achieves an AP of 56.4 on the OVCOD-D benchmark.

目标检测伪装物体开放词汇文本引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。