arXiv:2601.14255cs.CVcs.AI2026-01被引 3

用生成模型将粗略掩码转为精准视频抠像,零样本泛化能力强。

VideoMaMa: Mask-Guided Video Matting via Generative Prior

  • 基于预训练视频扩散模型,从粗掩码生成精细透明度图。
  • 在5万+真实视频上构建高质量抠像数据集,无需人工标注。
  • 适合做视频合成、特效制作的研究者和开发者使用。

由于标注数据稀缺,将视频抠像模型推广到真实场景仍具挑战。为此,我们提出视频掩码到抠像模型(VideoMaMa),通过利用预训练视频扩散模型,将粗略的分割掩码转换为像素级精确的透明度图。VideoMaMa 在仅用合成数据训练的情况下,仍展现出对真实视频的强零样本泛化能力。基于此,我们开发了可扩展的伪标签流水线,构建了「视频中任意物体抠像」(Matting Anything in Video, MA-V)数据集,包含超过50,000个真实世界视频的高质量抠像标注,覆盖多样场景与运动。为验证该数据集的有效性,我们在MA-V上微调SAM2模型,得到SAM2-Matte,其在真实视频上的鲁棒性优于在现有抠像数据集上训练的同类模型。这些结果凸显了大规模伪标注视频抠像的重要性,并展示了生成先验与可获取分割线索如何推动视频抠像研究的规模化进展。

原文摘要 · Abstract (English)

Generalizing video matting models to real-world videos remains a significant challenge due to the scarcity of labeled data. To address this, we present Video Mask-to-Matte Model (VideoMaMa) that converts coarse segmentation masks into pixel accurate alpha mattes, by leveraging pretrained video diffusion models. VideoMaMa demonstrates strong zero-shot generalization to real-world footage, even though it is trained solely on synthetic data. Building on this capability, we develop a scalable pseudo-labeling pipeline for large-scale video matting and construct the Matting Anything in Video (MA-V) dataset, which offers high-quality matting annotations for more than 50K real-world videos spanning diverse scenes and motions. To validate the effectiveness of this dataset, we fine-tune the SAM2 model on MA-V to obtain SAM2-Matte, which outperforms the same model trained on existing matting datasets in terms of robustness on in-the-wild videos. These findings emphasize the importance of large-scale pseudo-labeled video matting and showcase how generative priors and accessible segmentation cues can drive scalable progress in video matting research.

视频抠像生成模型伪标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。