构建4K级高精度语义分割数据集MaSS13K,推动图像编辑与AR/VR应用
MaSS13K: A Matting-level Semantic Segmentation Benchmark
- 提出多阶段特征融合的高效像素解码器,实现高分辨率掩码生成
- 数据集含1.3万张4K图像,掩码复杂度是现有数据集的20-50倍
- 新训练范式使模型可泛化到未标注类别,适合高精度视觉任务研究
高分辨率语义分割对图像编辑、虚实融合成像、AR/VR等应用至关重要。然而,现有数据集分辨率有限,且缺乏精确的掩码细节与边界。本文构建大规模、精细级语义分割数据集MaSS13K,包含13,348张真实世界4K分辨率图像,涵盖人体、植被、地面、天空、水体、建筑和其他七类对象,提供高质量掩码标注。其平均掩码复杂度为现有数据集的20-50倍。为此,提出专为高分辨率设计的MaSSFormer模型,采用三阶段特征融合的高效像素解码器,整合高层语义与低层纹理信息,以最小计算开销生成高分辨率掩码。进一步提出新学习范式,将七类高质量掩码与新类别伪标签结合,使MaSSFormer可迁移至其他物体类别。在MaSS13K上与14个代表性模型对比评估,验证了其优越性能。数据集与代码见https://github.com/xiechenxi99/MaSS13K。
原文摘要 · Abstract (English)
High-resolution semantic segmentation is essential for applications such as image editing, bokeh imaging, AR/VR, etc. Unfortunately, existing datasets often have limited resolution and lack precise mask details and boundaries. In this work, we build a large-scale, matting-level semantic segmentation dataset, named MaSS13K, which consists of 13,348 real-world images, all at 4K resolution. MaSS13K provides high-quality mask annotations of a number of objects, which are categorized into seven categories: human, vegetation, ground, sky, water, building, and others. MaSS13K features precise masks, with an average mask complexity 20-50 times higher than existing semantic segmentation datasets. We consequently present a method specifically designed for high-resolution semantic segmentation, namely MaSSFormer, which employs an efficient pixel decoder that aggregates high-level semantic features and low-level texture features across three stages, aiming to produce high-resolution masks with minimal computational cost. Finally, we propose a new learning paradigm, which integrates the high-quality masks of the seven given categories with pseudo labels from new classes, enabling MaSSFormer to transfer its accurate segmentation capability to other classes of objects. Our proposed MaSSFormer is comprehensively evaluated on the MaSS13K benchmark together with 14 representative segmentation models. We expect that our meticulously annotated MaSS13K dataset and the MaSSFormer model can facilitate the research of high-resolution and high-quality semantic segmentation. Datasets and codes can be found at https://github.com/xiechenxi99/MaSS13K.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。