提出8项指标的说话头视频评估框架,更全面衡量生成质量。
THEval. Evaluation Framework for Talking Head Video Generation
- 从质量、自然度、同步性三维度设计8项评估指标
- 17个顶尖模型在8.5万视频测试中显示同步好但表情与细节差
- 新数据集+开源代码/榜单,助力领域持续评测
视频生成已取得显著进展,生成内容日益逼真。然而,生成技术的快速发展远超评估指标的发展速度。当前说话头视频评估主要依赖有限指标,涵盖视频整体质量、口型同步性,并辅以用户研究。为此,我们提出一个包含8项指标的新评估框架,覆盖质量、自然度和同步性三个维度。选指标时注重效率与人类偏好的一致性,重点分析头部、口部和眉毛的细粒度动态变化及面部质量。在17个先进模型生成的8.5万条视频上进行广泛实验发现,尽管多数算法在口型同步方面表现良好,但在生成表情丰富性和无瑕疵细节方面仍面临挑战。这些视频基于我们构建的新真实数据集生成,以缓解训练数据偏差。所提出的基准框架旨在评估生成方法的改进效果。原始代码、数据集及排行榜将公开发布并定期更新,以反映领域最新进展。
原文摘要 · Abstract (English)
Video generation has achieved remarkable progress, with generated videos increasingly resembling real ones. However, the rapid advance in generation has outpaced the development of adequate evaluation metrics. Currently, the assessment of talking head generation primarily relies on limited metrics, evaluating general video quality, lip synchronization, and on conducting user studies. Motivated by this, we propose a new evaluation framework comprising 8 metrics related to three dimensions (i) quality, (ii) naturalness, and (iii) synchronization. In selecting the metrics, we place emphasis on efficiency, as well as alignment with human preferences. Based on this considerations, we streamline to analyze fine-grained dynamics of head, mouth, and eyebrows, as well as face quality. Our extensive experiments on 85,000 videos generated by 17 state-of-the-art models suggest that while many algorithms excel in lip synchronization, they face challenges with generating expressiveness and artifact-free details. These videos were generated based on a novel real dataset, that we have curated, in order to mitigate bias of training data. Our proposed benchmark framework is aimed at evaluating the improvement of generative methods. Original code, dataset and leaderboards will be publicly released and regularly updated with new methods, in order to reflect progress in the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。