arXiv:2510.21605cs.CV2025-10被引 6

用合成数据提升显著性物体检测泛化能力,无需真实标注

S3OD: Towards Generalizable Salient Object Detection with Synthetic Data

  • 通过多模态扩散模型生成超13万张高分辨率合成图像
  • 仅用合成数据训练的模型跨数据集误差降低20%-50%
  • 设计多掩码解码器应对检测歧义,适合跨任务迁移

显著性物体检测属于数据受限任务,昂贵的像素级标注导致不同子任务(如DIS和HR-SOD)需独立训练。本文提出S3OD方法,通过大规模合成数据生成与模糊感知架构显著提升泛化能力。构建了超过13.9万张高分辨率图像的S3OD数据集,基于多模态扩散流水线从扩散模型和DINO-v3特征中提取标签。迭代生成框架根据模型表现优先生成挑战类别。提出轻量级多掩码解码器,通过预测多个有效解释来处理显著性检测中的固有歧义。仅在合成数据上训练的模型在跨数据集测试中实现20%-50%的误差降低,微调后版本在DIS和HR-SOD基准上达到当前最优性能。

原文摘要 · Abstract (English)

Salient object detection exemplifies data-bounded tasks where expensive pixel-precise annotations force separate model training for related subtasks like DIS and HR-SOD. We present a method that dramatically improves generalization through large-scale synthetic data generation and ambiguity-aware architecture. We introduce S3OD, a dataset of over 139,000 high-resolution images created through our multi-modal diffusion pipeline that extracts labels from diffusion and DINO-v3 features. The iterative generation framework prioritizes challenging categories based on model performance. We propose a streamlined multi-mask decoder that handles the inherent ambiguity in salient object detection by predicting multiple valid interpretations. Models trained only on synthetic data achieve 20-50% error reduction in cross-dataset generalization, while fine-tuned versions reach state-of-the-art performance across DIS and HR-SOD benchmarks.

显著性检测合成数据多掩码泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。