arXiv:2501.06038cs.CV2025-01被引 1

用点标注+文本引导,弱监督检测伪装物体效果更优

A Holistically Point-guided Text Framework for Weakly-Supervised Camouflaged Object Detection

  • 分三阶段:生成候选掩码、筛选优质掩码、自监督训练
  • 在四个基准数据集上显著超越现有方法,部分超全监督模型
  • 提出新数据集P2C-COD和T-COD,推动弱监督伪装目标检测发展

弱监督伪装目标检测(WSCOD)通过弱标签训练模型以分割与背景视觉融合的物体,近年来备受关注。尽管稀疏标注(如涂鸦)已取得良好效果,但点-文本监督仍鲜有研究。本文提出一种全新的整体式点引导文本框架,分为三个阶段:生成、选择、训练。首先设计点引导候选生成(PCG),利用点标注的前景信息修正文本路径,提升掩码生成质量;其次引入合格候选判别器(QCD),基于CLIP从文本提示中筛选最优掩码;最后使用选定伪掩码进行自监督视觉变换器训练。同时构建了新的点监督数据集P2C-COD和文本监督数据集T-COD。在四个基准数据集上的实验证明,本方法显著优于当前最先进方法,部分甚至超过现有全监督方法。

原文摘要 · Abstract (English)

Weakly-Supervised Camouflaged Object Detection (WSCOD) has gained popularity for its promise to train models with weak labels to segment objects that visually blend into their surroundings. Recently, some methods using sparsely-annotated supervision shown promising results through scribbling in WSCOD, while point-text supervision remains underexplored. Hence, this paper introduces a novel holistically point-guided text framework for WSCOD by decomposing into three phases: segment, choose, train. Specifically, we propose Point-guided Candidate Generation (PCG), where the point's foreground serves as a correction for the text path to explicitly correct and rejuvenate the loss detection object during the mask generation process (SEGMENT). We also introduce a Qualified Candidate Discriminator (QCD) to choose the optimal mask from a given text prompt using CLIP (CHOOSE), and employ the chosen pseudo mask for training with a self-supervised Vision Transformer (TRAIN). Additionally, we developed a new point-supervised dataset (P2C-COD) and a text-supervised dataset (T-COD). Comprehensive experiments on four benchmark datasets demonstrate our method outperforms state-of-the-art methods by a large margin, and also outperforms some existing fully-supervised camouflaged object detection methods.

弱监督伪装检测文本引导点标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。