arXiv:2601.12080cs.CV2026-01AAAI

提升真实场景下图像抠图与分割精度,支持交互式精准定位。

Toward Real-World High-Precision Image Matting and Segmentation

  • 通过深度感知知识蒸馏增强前景表征能力。
  • 采用领域不变学习策略,提升合成数据在真实场景的泛化性能。
  • 支持视觉与语言双模态提示,实现灵活交互式预测。

高精度场景解析任务(如图像抠图与二值分割)旨在精确预测包含细粒度细节(如头发)的掩码。现有方法多聚焦于显著的单个前景物体,交互式方法虽可调整目标,但其类别无关设计限制了跨类泛化能力。此外,高质量标注数据稀缺导致依赖不一致的合成数据,造成真实场景泛化性能差。为此,我们提出前景一致性学习模型FCLM。首先引入深度感知知识蒸馏,以更好表征前景;针对数据困境,将合成数据处理视为领域自适应问题,提出领域不变学习策略,专注前景学习。为支持交互预测,设计基于对象的解码器,可接收视觉与语言提示,准确生成参考目标。实验表明,该方法在定量与定性上均优于当前最优方法。

原文摘要 · Abstract (English)

High-precision scene parsing tasks, including image matting and dichotomous segmentation, aim to accurately predict masks with extremely fine details (such as hair). Most existing methods focus on salient, single foreground objects. While interactive methods allow for target adjustment, their class-agnostic design restricts generalization across different categories. Furthermore, the scarcity of high-quality annotation has led to a reliance on inharmonious synthetic data, resulting in poor generalization to real-world scenarios. To this end, we propose a Foreground Consistent Learning model, dubbed as FCLM, to address the aforementioned issues. Specifically, we first introduce a Depth-Aware Distillation strategy where we transfer the depth-related knowledge for better foreground representation. Considering the data dilemma, we term the processing of synthetic data as domain adaptation problem where we propose a domain-invariant learning strategy to focus on foreground learning. To support interactive prediction, we contribute an Object-Oriented Decoder that can receive both visual and language prompts to predict the referring target. Experimental results show that our method quantitatively and qualitatively outperforms SOTA methods.

图像抠图交互式分割领域自适应多模态提示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。