arXiv:2501.04582cs.CV2025-01

用大模型生成精准伪标签,提升显著物体检测效果

Boosting Salient Object Detection with Knowledge Distillated from Large Foundation Models

  • 用文本提示引导大模型生成伪标签,降低标注成本
  • 新数据集BDS-TR规模更大、场景更丰富,推动研究发展
  • 动态上采样边解码器增强边缘细节,适合图像分割任务

显著物体检测(SOD)旨在识别并分割图像中的突出区域。传统方法依赖人工标注的精确像素级伪标签,耗时费力。本文提出一种低成本高精度的标注方法,利用大基础模型通过文本提示生成伪标签。由于大模型对图像显著区域关注不足,我们手动标注部分文本以微调模型。基于此方法,实现了快速且精准的伪标签生成,并构建了新数据集BDS-TR。相比之前的DUTS-TR,BDS-TR在规模、类别多样性和场景覆盖上均有提升,有助于增强模型在多场景下的适用性,并为未来SOD研究提供更全面的基础数据集。此外,本文提出基于动态上采样的边缘解码器,专注恢复物体边缘同时逐步提升特征分辨率。在五个基准数据集上的全面实验表明,该方法显著优于现有最先进方法,甚至超越多个全监督SOD方法。代码与结果将公开。

原文摘要 · Abstract (English)

Salient Object Detection (SOD) aims to identify and segment prominent regions within a scene. Traditional models rely on manually annotated pseudo labels with precise pixel-level accuracy, which is time-consuming. We developed a low-cost, high-precision annotation method by leveraging large foundation models to address the challenges. Specifically, we use a weakly supervised approach to guide large models in generating pseudo-labels through textual prompts. Since large models do not effectively focus on the salient regions of images, we manually annotate a subset of text to fine-tune the model. Based on this approach, which enables precise and rapid generation of pseudo-labels, we introduce a new dataset, BDS-TR. Compared to the previous DUTS-TR dataset, BDS-TR is more prominent in scale and encompasses a wider variety of categories and scenes. This expansion will enhance our model's applicability across a broader range of scenarios and provide a more comprehensive foundational dataset for future SOD research. Additionally, we present an edge decoder based on dynamic upsampling, which focuses on object edges while gradually recovering image feature resolution. Comprehensive experiments on five benchmark datasets demonstrate that our method significantly outperforms state-of-the-art approaches and also surpasses several existing fully-supervised SOD methods. The code and results will be made available.

显著物体检测伪标签大模型边缘优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。