构建大规模多样化的音视频数据集,提升语音驱动人脸生成的泛化能力
TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis
- 基于多阶段自动化筛选,构建1244小时、7729人参与的高质量数据集
- 在跨数据集测试中,基于该数据训练的模型表现更优,泛化能力显著提升
- 提供分层评估集,揭示传统指标忽略的群体性能差异,适合公平性研究
语音驱动的人脸生成已实现惊人的真实感,但现有最先进模型在族裔、语言和年龄等人类多样性方面缺乏泛化能力。我们认为,这一差距源于训练数据的规模、质量和多样性不足。为此,我们推出TalkVid,一个大规模、高质量、多样化的数据集,包含来自7729位独特说话人的1244小时视频。该数据集通过多阶段自动化流程精心构建,严格过滤运动稳定性、美学质量与面部细节,并经人工评估验证可靠性。此外,我们构建并发布TalkVid-Bench,一个500段视频的分层评估集,按关键人口与语言维度均衡分布。实验表明,基于TalkVid训练的模型在跨数据集泛化上优于以往模型。关键的是,对TalkVid-Bench的分析揭示了传统聚合指标掩盖的子群体性能差异,凸显其对后续研究的重要性。代码与数据见 https://github.com/FreedomIntelligence/TalkVid
原文摘要 · Abstract (English)
Audio-driven talking head synthesis has achieved remarkable photorealism, yet state-of-the-art (SOTA) models exhibit a critical failure: they lack generalization to the full spectrum of human diversity in ethnicity, language, and age groups. We argue that this generalization gap is a direct symptom of limitations in existing training data, which lack the necessary scale, quality, and diversity. To address this challenge, we introduce TalkVid, a new large-scale, high-quality, and diverse dataset containing 1244 hours of video from 7729 unique speakers. TalkVid is curated through a principled, multi-stage automated pipeline that rigorously filters for motion stability, aesthetic quality, and facial detail, and is validated against human judgments to ensure its reliability. Furthermore, we construct and release TalkVid-Bench, a stratified evaluation set of 500 clips meticulously balanced across key demographic and linguistic axes. Our experiments demonstrate that a model trained on TalkVid outperforms counterparts trained on previous datasets, exhibiting superior cross-dataset generalization. Crucially, our analysis on TalkVid-Bench reveals performance disparities across subgroups that are obscured by traditional aggregate metrics, underscoring its necessity for future research. Code and data can be found in https://github.com/FreedomIntelligence/TalkVid
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。