arXiv:2512.01803cs.CV2025-12

用真实人体动作数据训练新指标,更准评估生成视频中的人体动作合理性。

Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos

  • 融合骨骼结构与视觉特征,构建无外观依赖的动作表征空间。
  • 在自建基准上比现有方法提升超68%,与人眼判断更吻合。
  • 适合研究视频生成、动作质量评估的学者和工程师使用。

尽管视频生成模型进展迅速,但对复杂人体动作的视觉与时间正确性仍缺乏可靠评估指标。现有纯视觉编码器和多模态大语言模型严重依赖外观,缺乏时间理解能力,难以识别生成视频中细微的动作动态与解剖不合理之处。本文提出一种基于真实人体动作学习的潜在空间评估指标,通过融合无外观依赖的人体骨骼几何特征与视觉特征,捕捉真实动作的细微约束与时间平滑性。该联合特征空间提供动作合理性的稳健表示。给定生成视频后,通过计算其潜在表示与真实动作分布之间的距离来量化动作质量。为严格验证,我们构建了一个专门探测时间挑战性的人体动作保真度的多维度新基准。大量实验表明,该指标在新基准上相比现有最先进方法提升超过68%,在外部基准上表现良好,且与人类感知相关性更强。深入分析揭示了当前视频生成模型的关键缺陷,确立了视频生成高级研究的新标准。

原文摘要 · Abstract (English)

Despite rapid advances in video generative models, robust metrics for evaluating visual and temporal correctness of complex human actions remain elusive. Critically, existing pure-vision encoders and Multimodal Large Language Models (MLLMs) are strongly appearance-biased, lack temporal understanding, and thus struggle to discern intricate motion dynamics and anatomical implausibilities in generated videos. We tackle this gap by introducing a novel evaluation metric derived from a learned latent space of real-world human actions. Our method first captures the nuances, constraints, and temporal smoothness of real-world motion by fusing appearance-agnostic human skeletal geometry features with appearance-based features. We posit that this combined feature space provides a robust representation of action plausibility. Given a generated video, our metric quantifies its action quality by measuring the distance between its underlying representations and this learned real-world action distribution. For rigorous validation, we develop a new multi-faceted benchmark specifically designed to probe temporally challenging aspects of human action fidelity. Through extensive experiments, we show that our metric achieves substantial improvement of more than 68% compared to existing state-of-the-art methods on our benchmark, performs competitively on established external benchmarks, and has a stronger correlation with human perception. Our in-depth analysis reveals critical limitations in current video generative models and establishes a new standard for advanced research in video generation.

动作评估视频生成生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。