通过迭代优化提升真实场景图像抠图精度。
Enhancing Image Matting in Real-World Scenes with Mask-Guided Iterative Refinement
- 用自监督视觉变换器特征增强语义理解。
- 在多个真实数据集上达到当前最佳性能。
- 适合需要高精度抠图的影视与AR应用。
真实世界图像抠图在内容创作和增强现实中有重要应用,但因场景复杂且高质量数据集稀缺而面临挑战。为此,我们提出Mask2Alpha,一种迭代精炼框架,旨在提升语义理解、实例区分与细节恢复能力。该框架利用自监督视觉变换器特征作为语义先验,增强复杂场景中的上下文理解;通过掩码引导的特征选择模块,实现多实例场景下对目标对象的精准定位;并采用基于稀疏卷积的优化方案,从低分辨率语义传递逐步细化至高分辨率稀疏重建,有效恢复精细细节。在多个真实世界数据集上的基准测试显示,Mask2Alpha持续取得最优结果,验证了其在准确性和效率上的卓越表现。
原文摘要 · Abstract (English)
Real-world image matting is essential for applications in content creation and augmented reality. However, it remains challenging due to the complex nature of scenes and the scarcity of high-quality datasets. To address these limitations, we introduce Mask2Alpha, an iterative refinement framework designed to enhance semantic comprehension, instance awareness, and fine-detail recovery in image matting. Our framework leverages self-supervised Vision Transformer features as semantic priors, strengthening contextual understanding in complex scenarios. To further improve instance differentiation, we implement a mask-guided feature selection module, enabling precise targeting of objects in multi-instance settings. Additionally, a sparse convolution-based optimization scheme allows Mask2Alpha to recover high-resolution details through progressive refinement,from low-resolution semantic passes to high-resolution sparse reconstructions. Benchmarking across various real-world datasets, Mask2Alpha consistently achieves state-of-the-art results, showcasing its effectiveness in accurate and efficient image matting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。