arXiv:2512.20260cs.CVcs.AI2025-12被引 3

用辩论机制提升弱监督下的伪装目标检测精度

Debate-Enhanced Pseudo Labeling and Frequency-Aware Progressive Debiasing for Weakly-Supervised Camouflaged Object Detection with Scribble Annotations

  • 通过多智能体辩论与自适应采样优化伪标签生成
  • 提出频域感知渐进去偏网络,缓解涂鸦标注偏差
  • 适合研究弱监督、伪装目标检测的开发者

弱监督伪装目标检测(WSCOD)旨在仅依赖稀疏涂鸦标注的情况下定位并分割与背景视觉融合的目标。现有方法仍远落后于全监督模型,主要因两点:一是通用分割模型(如SAM)生成的伪掩码不可靠,缺乏任务特定语义理解;二是忽略涂鸦标注固有的偏差,难以捕捉伪装目标的全局结构。为此,本文提出两阶段框架${D}^{3}$ETOR,第一阶段引入自适应熵驱动点采样和多智能体辩论机制,增强SAM在伪装目标检测中的伪标签生成能力,提升解释性与精度;第二阶段设计FADeNet,通过逐步融合多层级频域感知特征,平衡全局语义与局部细节建模,并动态重加权不同区域的监督强度以缓解标注偏差。联合利用伪标签与涂鸦语义信号,${D}^{3}$ETOR显著缩小了弱监督与全监督之间的差距,在多个基准上达到当前最优性能。

原文摘要 · Abstract (English)

Weakly-Supervised Camouflaged Object Detection (WSCOD) aims to locate and segment objects that are visually concealed within their surrounding scenes, relying solely on sparse supervision such as scribble annotations. Despite recent progress, existing WSCOD methods still lag far behind fully supervised ones due to two major limitations: (1) the pseudo masks generated by general-purpose segmentation models (e.g., SAM) and filtered via rules are often unreliable, as these models lack the task-specific semantic understanding required for effective pseudo labeling in COD; and (2) the neglect of inherent annotation bias in scribbles, which hinders the model from capturing the global structure of camouflaged objects. To overcome these challenges, we propose ${D}^{3}$ETOR, a two-stage WSCOD framework consisting of Debate-Enhanced Pseudo Labeling and Frequency-Aware Progressive Debiasing. In the first stage, we introduce an adaptive entropy-driven point sampling method and a multi-agent debate mechanism to enhance the capability of SAM for COD, improving the interpretability and precision of pseudo masks. In the second stage, we design FADeNet, which progressively fuses multi-level frequency-aware features to balance global semantic understanding with local detail modeling, while dynamically reweighting supervision strength across regions to alleviate scribble bias. By jointly exploiting the supervision signals from both the pseudo masks and scribble semantics, ${D}^{3}$ETOR significantly narrows the gap between weakly and fully supervised COD, achieving state-of-the-art performance on multiple benchmarks.

伪装检测弱监督伪标签去偏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。