arXiv:2501.16147cs.CV2025-01

用文本提示和层扩散生成高质量人像抠图,再通过连通性先验优化结果。

Efficient Portrait Matte Creation With Layer Diffusion and Connectivity Priors

  • 结合文本提示与层扩散模型生成人像前景并提取隐式抠图
  • 构建包含20051张图像的高质量人像抠图数据集LD-Portrait-20K
  • 适用于人像视频抠图与图像抠图任务,提升现有方法性能

高效构建人像抠图模型需要大量高质量训练数据,但真实场景下难以获取。由于最精确的人像抠图需在绿幕前拍摄,难以大规模收集真实数据。本文提出利用文本提示与最新的层扩散模型生成高质量人像前景,并提取潜在抠图。然而,生成结果存在显著伪影。受人像图像中前景边界始终连通的先验启发,引入连通性感知方法对抠图进行优化。基于此,构建了大规模人像抠图数据集LD-Portrait-20K,包含20,051个人像前景及高精度α抠图。大量实验表明,基于该数据集训练的模型显著优于其他数据集上的模型。与色键算法对比及消融实验进一步验证了该方法的有效性。此外,该数据集还推动了基于图像抠图模型与简单视频分割的前沿视频人像抠图性能提升。

原文摘要 · Abstract (English)

Learning effective deep portrait matting models requires training data of both high quality and large quantity. Neither quality nor quantity can be easily met for portrait matting, however. Since the most accurate ground-truth portrait mattes are acquired in front of the green screen, it is almost impossible to harvest a large-scale portrait matting dataset in reality. This work shows that one can leverage text prompts and the recent Layer Diffusion model to generate high-quality portrait foregrounds and extract latent portrait mattes. However, the portrait mattes cannot be readily in use due to significant generation artifacts. Inspired by the connectivity priors observed in portrait images, that is, the border of portrait foregrounds always appears connected, a connectivity-aware approach is introduced to refine portrait mattes. Building on this, a large-scale portrait matting dataset is created, termed LD-Portrait-20K, with $20,051$ portrait foregrounds and high-quality alpha mattes. Extensive experiments demonstrated the value of the LD-Portrait-20K dataset, with models trained on it significantly outperforming those trained on other datasets. In addition, comparisons with the chroma keying algorithm and an ablation study on dataset capacity further confirmed the effectiveness of the proposed matte creation approach. Further, the dataset also contributes to state-of-the-art video portrait matting, implemented by simple video segmentation and a trimap-based image matting model trained on this dataset.

人像抠图扩散模型数据集视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。