用手绘草图实现精准图像分割,无需标注掩码。
SketchYourSeg: Mask-Free Subjective Image Segmentation via Freehand Sketches
- 以手绘草图为查询,直接实现跨图库的主观分割。
- 在多个基准上优于现有方法,支持多粒度分割。
- 无需像素级标注,适合普通用户快速操作。
我们提出SketchYourSeg,一种新框架,通过单张手绘草图实现整个图库中主观图像分割。与难以表达空间细节的文本提示或仅限单图交互的方法不同,草图天然融合语义意图与结构精度,可有效区分视觉相似实例、精确定位部件边界或表达组合概念的空间关系。该方法解决三大挑战:(i) 采用无掩码训练框架,避免像素级标注;(ii) 构建草图检索(SBIR)模型与基础模型(CLIP/DINOv2)间的协同机制,前者提供训练信号,后者生成掩码;(iii) 通过专用草图增强策略实现多粒度分割能力。大量实验表明,该方法在多个基准上表现优异,确立了一种兼顾精度与效率的用户引导分割新范式。
原文摘要 · Abstract (English)
We introduce SketchYourSeg, a novel framework that establishes freehand sketches as a powerful query modality for subjective image segmentation across entire galleries through a single exemplar sketch. Unlike text prompts that struggle with spatial specificity or interactive methods confined to single-image operations, sketches naturally combine semantic intent with structural precision. This unique dual encoding enables precise visual disambiguation for segmentation tasks where text descriptions would be cumbersome or ambiguous -- such as distinguishing between visually similar instances, specifying exact part boundaries, or indicating spatial relationships in composed concepts. Our approach addresses three fundamental challenges: (i) eliminating the need for pixel-perfect annotation masks during training with a mask-free framework; (ii) creating a synergistic relationship between sketch-based image retrieval (SBIR) models and foundation models (CLIP/DINOv2) where the former provides training signals while the latter generates masks; and (iii) enabling multi-granular segmentation capabilities through purpose-made sketch augmentation strategies. Our extensive evaluations demonstrate superior performance over existing approaches across diverse benchmarks, establishing a new paradigm for user-guided image segmentation that balances precision with efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。