arXiv:2512.11782cs.CV2025-12中稿 · CVPR被引 6

用可学习评估器提升视频抠图质量与数据规模

MatAnyone 2: Scaling Video Matting via a Learned Quality Evaluator

  • 引入可学习的质量评估器,无需真值即可评估抠图语义与边界质量
  • 构建28,000段、240万帧的实时视频抠图数据集VMReal
  • 适合需要高质量视频抠图和大规模数据训练的研究者

视频抠图受限于现有数据集的规模与真实感。尽管利用分割数据可提升语义稳定性,但缺乏有效边界监督常导致结果呈现类似分割的粗糙边缘。为此,我们提出一种可学习的抠图质量评估器(MQE),可在无真值情况下评估alpha抠图的语义与边界质量,生成像素级评估图以识别可靠与错误区域,实现细粒度质量评估。该方法通过两种方式扩展视频抠图:(1) 在训练中作为在线质量反馈,抑制错误区域,提供全面监督;(2) 作为离线数据筛选模块,结合领先视频与图像抠图模型优势,提升标注质量。由此构建了包含28,000段视频、240万帧的大规模真实世界视频抠图数据集VMReal。为应对长视频中的显著外观变化,引入参考帧训练策略,将局部窗口外的远距离帧纳入训练。所提出的MatAnyone 2在合成与真实世界基准上均达到最先进性能,各项指标全面超越先前方法。

原文摘要 · Abstract (English)

Video matting remains limited by the scale and realism of existing datasets. While leveraging segmentation data can enhance semantic stability, the lack of effective boundary supervision often leads to segmentation-like mattes lacking fine details. To this end, we introduce a learned Matting Quality Evaluator (MQE) that assesses semantic and boundary quality of alpha mattes without ground truth. It produces a pixel-wise evaluation map that identifies reliable and erroneous regions, enabling fine-grained quality assessment. The MQE scales up video matting in two ways: (1) as an online matting-quality feedback during training to suppress erroneous regions, providing comprehensive supervision, and (2) as an offline selection module for data curation, improving annotation quality by combining the strengths of leading video and image matting models. This process allows us to build a large-scale real-world video matting dataset, VMReal, containing 28K clips and 2.4M frames. To handle large appearance variations in long videos, we introduce a reference-frame training strategy that incorporates long-range frames beyond the local window for effective training. Our MatAnyone 2 achieves state-of-the-art performance on both synthetic and real-world benchmarks, surpassing prior methods across all metrics.

视频抠图数据集构建质量评估训练策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。