arXiv:2410.07151cs.CV2024-10ICCV被引 14

构建1200小时高质量人脸视频数据集,解决亚洲面孔代表性不足问题

DH-FaceVid-1K: A Large-Scale High-Quality Dataset for Face Video Generation

  • 收集2万+人、27万段视频,含语音与面部关键点标注
  • 涵盖1200小时视频,支持文本/图像生成人脸视频任务
  • 重点补足亚洲面孔数据,降低模型偏见

以人为中心的生成模型日益流行,催生了基于文本或音频提示生成说话人脸视频等创新应用。其核心依赖于在大规模高质量数据上预训练的基础模型。然而,许多先进方法依赖受限制的内部数据,现有公开数据集普遍缺乏高分辨率人脸视频。本文提出大规模人脸视频数据集DH-FaceVid-1K,总时长1200小时,包含270,043段视频片段,来自超过20,000名个体。每条样本均配有对应语音音频、面部关键点及文本标注。相比现有公开数据集,本数据集具备多民族覆盖和高质量、全面的个体属性信息。我们建立了多个支持文本到视频、图像到视频生成的人脸视频生成模型,并构建了完整基准评估不同数据比例下的缩放规律。主要目标是弥补现有数据集中亚洲面孔的代表性不足,丰富全球以人为核心的语料库,缓解人口统计偏差。

原文摘要 · Abstract (English)

Human-centric generative models are becoming increasingly popular, giving rise to various innovative tools and applications, such as talking face videos conditioned on text or audio prompts. The core of these capabilities lies in powerful pre-trained foundation models, trained on large-scale, high-quality datasets. However, many advanced methods rely on in-house data subject to various constraints, and other current studies fail to generate high-resolution face videos, which is mainly attributed to the significant lack of large-scale, high-quality face video datasets. In this paper, we introduce a human face video dataset, \textbf{DH-FaceVid-1K}. Our collection spans 1,200 hours in total, encompassing 270,043 video clips from over 20,000 individuals. Each sample includes corresponding speech audio, facial keypoints, and text annotations. Compared to other publicly available datasets, ours distinguishes itself through its multi-ethnic coverage and high-quality, comprehensive individual attributes. We establish multiple face video generation models supporting tasks such as text-to-video and image-to-video generation. In addition, we develop comprehensive benchmarks to validate the scaling law when using different proportions of proposed dataset. Our primary aim is to contribute a face video dataset, particularly addressing the underrepresentation of Asian faces in existing curated datasets and thereby enriching the global spectrum of face-centric data and mitigating demographic biases. \textbf{Project Page:} https://luna-ai-lab.github.io/DH-FaceVid-1K/

人脸生成数据集多民族覆盖

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。