构建首个万级对话头质量评估数据集,评测AI生成视频优劣。
Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads
- 构建10457个AI生成说话头视频的数据集,涵盖12个文生图模型与14种驱动方法。
- 发现现有生成结果普遍存在口型与语音不同步、表情失真等质量问题。
- 提出基于首帧、音调-唇动一致性的客观评估方法,性能达当前最优。
以语音驱动的人像生成技术被称为“说话头”(Talkers),随着文本到图像(T2I)模型的快速发展,人工智能生成的说话头(AGTHs)逐渐成为新兴数字人媒体。然而,其生成质量仍面临挑战,且相关系统性研究较少。为此,本文构建了迄今最大的AGTH质量评估数据集THQA-10K,选取12个主流T2I模型和14种先进说话头方法,针对14个提示词生成视频。剔除失败样本后,共获得10,457个有效视频。通过志愿者主观评分并标注失真类别,分析显示现有方法在泛化性和质量上存在明显不足,并揭示常见失真模式。最后提出一种基于首帧、音调-唇动一致性(Y-T slice)的客观评估方法,实验表明该方法在AGTH质量评估任务中达到当前最佳性能。项目代码已开源:https://github.com/zyj-2000/Talker。
原文摘要 · Abstract (English)
Speech-driven methods for portraits are figuratively known as "Talkers" because of their capability to synthesize speaking mouth shapes and facial movements. Especially with the rapid development of the Text-to-Image (T2I) models, AI-Generated Talking Heads (AGTHs) have gradually become an emerging digital human media. However, challenges persist regarding the quality of these talkers and AGTHs they generate, and comprehensive studies addressing these issues remain limited. To address this gap, this paper presents the largest AGTH quality assessment dataset THQA-10K to date, which selects 12 prominent T2I models and 14 advanced talkers to generate AGTHs for 14 prompts. After excluding instances where AGTH generation is unsuccessful, the THQA-10K dataset contains 10,457 AGTHs. Then, volunteers are recruited to subjectively rate the AGTHs and give the corresponding distortion categories. In our analysis for subjective experimental results, we evaluate the performance of talkers in terms of generalizability and quality, and also expose the distortions of existing AGTHs. Finally, an objective quality assessment method based on the first frame, Y-T slice and tone-lip consistency is proposed. Experimental results show that this method can achieve state-of-the-art (SOTA) performance in AGTH quality assessment. The work is released at https://github.com/zyj-2000/Talker.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。