首个系统评估生成视频中人体动作真实性的基准框架
HumanScore: Benchmarking Human Motions in Generated Videos

- 设计六项可解释指标,涵盖运动学合理性与生物力学一致性
- 13个主流模型对比显示:视觉真实度高但动作生物力学存在普遍缺陷
- 适合关注视频生成质量评估、动作仿真研究的学者与开发者
近期模型架构、算力和数据规模的进步推动了视频生成技术的快速发展,产出越来越逼真的内容。然而,此前缺乏系统性方法来衡量这些系统对人体姿态与运动动态的还原程度。本文提出HumanScore,一个系统性评估生成视频中人体动作质量的框架。该框架定义了六项可解释的指标,覆盖运动学合理性、时间稳定性与生物力学一致性,支持细粒度诊断,超越单纯视觉真实性的评价。通过精心设计的提示词,我们诱发多样化的动作,涵盖不同强度,并评估十三个前沿生成模型的表现。分析揭示感知合理性与动作生物力学保真度之间存在持续差距,识别出常见失败模式(如时间抖动、解剖上不可能的姿态、运动漂移),并基于量化且具有物理意义的标准,建立可靠的模型排名。
原文摘要 · Abstract (English)
Recent advances in model architectures, compute, and data scale have driven rapid progress in video generation, producing increasingly realistic content. Yet, no prior method systematically measures how faithfully these systems render human bodies and motion dynamics. In this paper, we present HumanScore, a systematic framework to evaluate the quality of human motions in AI-generated videos. HumanScore defines six interpretable metrics spanning kinematic plausibility, temporal stability, and biomechanical consistency, enabling fine-grained diagnosis beyond visual realism alone. Through carefully designed prompts, we elicit a diverse set of movements at varying intensities and evaluate videos generated by thirteen state-of-the-art models. Our analysis reveals consistent gaps between perceptual plausibility and motion biomechanical fidelity, identifies recurrent failure modes (e.g., temporal jitter, anatomically implausible poses, and motion drift), and produces robust model rankings from quantitative and physically meaningful criteria.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。