统一支持多种交互方式的高精度图像分割模型
FocalClick-XL: Towards Unified and High-quality Interactive Segmentation
- 分层级设计三子网,分别处理上下文、物体和细节信息
- 在点击/涂鸦/框等交互下均达当前最佳性能
- 可生成带精细边缘的透明度图,适合多场景应用
交互式分割允许用户通过点击、涂鸦或框选等方式提取目标对象的二值掩码。然而现有方法通常仅支持有限的交互形式,且难以捕捉细粒度细节。本文重新审视FocalClick的经典粗到精设计,提出新框架FocalClick-XL。受大规模预训练趋势启发,将交互分割分解为上下文、物体和细节三个元任务,分别为每个层次分配专用子网。这种分解使各子网可独立进行大规模预训练,提升效果。为增强灵活性,在不同交互形式间共享上下文与细节信息作为通用知识,并在物体层级引入提示层以编码特定交互类型。结果表明,FocalClick-XL在基于点击的基准上达到最先进性能,且对框、涂鸦、粗略掩码等多种交互格式表现出卓越适应性。此外,该模型还能生成包含精细边缘的α蒙版,成为功能强大的交互式分割工具。
原文摘要 · Abstract (English)
Interactive segmentation enables users to extract binary masks of target objects through simple interactions such as clicks, scribbles, and boxes. However, existing methods often support only limited interaction forms and struggle to capture fine details. In this paper, we revisit the classical coarse-to-fine design of FocalClick and introduce significant extensions. Inspired by its multi-stage strategy, we propose a novel pipeline, FocalClick-XL, to address these challenges simultaneously. Following the emerging trend of large-scale pretraining, we decompose interactive segmentation into meta-tasks that capture different levels of information -- context, object, and detail -- assigning a dedicated subnet to each level.This decomposition allows each subnet to undergo scaled pretraining with independent data and supervision, maximizing its effectiveness. To enhance flexibility, we share context- and detail-level information across different interaction forms as common knowledge while introducing a prompting layer at the object level to encode specific interaction types. As a result, FocalClick-XL achieves state-of-the-art performance on click-based benchmarks and demonstrates remarkable adaptability to diverse interaction formats, including boxes, scribbles, and coarse masks. Beyond binary mask generation, it is also capable of predicting alpha mattes with fine-grained details, making it a versatile and powerful tool for interactive segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。