arXiv:2604.25361cs.CV2026-04中稿 · the 2026 IEEE Inte…

提出细粒度人体视频评估框架,更贴近人眼判断标准。

HuM-Eval: A Coarse-to-Fine Framework for Human-Centric Video Evaluation

论文配图:HuM-Eval: A Coarse-to-Fine Framework for Human-Centric Video Evaluation
图 1 · 摘自论文原文
  • 分粗到细两阶段评估:先看整体画面,再检人体姿态与运动稳定性。
  • 在1000个提示上测试,平均人评相关性达58.2%,优于现有方法。
  • 适合研究视频生成、人体动作建模的学者和开发者使用。

近年来视频生成模型快速发展,其中自然人体动作生成至关重要。然而,准确评估生成人体动作视频的质量仍面临挑战。现有评估指标多关注全局场景统计特征,忽略细粒度人体细节,导致与人类主观偏好不一致。为此,我们提出HuM-Eval,一种新型以人为中心的评估框架,采用粗到细策略:首先利用视觉语言模型进行全局视频质量粗评;随后通过2D姿态验证解剖结构正确性,3D人体运动评估动作稳定性。大量实验表明,HuM-Eval平均人评相关性达58.2%,超越现有最优基线。此外,我们构建了涵盖1000个多样化提示的HuM-Bench基准,对现有文本到视频模型进行详细评估,为下一代人体动作生成提供支持。

原文摘要 · Abstract (English)

Video generation models have developed rapidly in recent years, where generating natural human motion plays a pivotal role. However, accurately evaluating the quality of generated human motion video remains a significant challenge. Existing evaluation metrics primarily focus on global scene statistics, often overlooking fine-grained human details and consequently failing to align with human subjective preference. To bridge this gap, we propose HuM-Eval, a novel human-centric evaluation framework that adopts a coarse-to-fine strategy. Specifically, our framework first utilizes a Vision Language Model to perform a coarse assessment of global video quality. It then proceeds to a fine-grained analysis, using 2D pose to verify anatomical correctness and 3D human motion to evaluate motion stability. Extensive experiments demonstrate that HuM-Eval achieves an average human correlation of 58.2%, outperforming state-of-the-art baselines. Furthermore, we introduce HuM-Bench, a comprehensive benchmark comprising 1,000 diverse prompts, and conduct a detailed evaluation of existing text-to-video models, paving the way for next-generation human motion generation.

视频评估人体动作多模态基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。