arXiv:2503.04666cs.CV2025-03被引 4

构建首个细粒度人类视频生成评估基准,助力模型能力精准衡量。

What Are You Doing? A Closer Look at Controllable Human Video Generation

  • 设计包含56类动作的1544个带标注视频数据集
  • 提出自动评估指标,与人工评价高度一致
  • 首次系统分析7个主流模型在动作、交互等9维度表现

高质量基准对推动机器学习研究至关重要。尽管视频生成领域日益热门,但缺乏全面的人类视频生成评估数据集。现有数据集如TikTok和TED-Talks在动作多样性与复杂性上不足,难以充分检验生成模型能力。为此,我们推出「你在做什么?」(WYD):一个用于可控图像到视频生成中人类行为细粒度评估的新基准。WYD包含1,544个带标注视频,涵盖56种细粒度类别,可系统评估生成模型在动作、交互与运动等9个方面的表现。我们还提出并验证了基于标注信息的自动评估指标,能更好反映人工评价结果。基于该数据集与指标,我们深入分析了7个前沿可控图像到视频生成模型,揭示其性能差异与局限。相关数据与代码已开源,网址:https://github.com/google-deepmind/wyd-benchmark。

原文摘要 · Abstract (English)

High-quality benchmarks are crucial for driving progress in machine learning research. However, despite the growing interest in video generation, there is no comprehensive dataset to evaluate human generation. Humans can perform a wide variety of actions and interactions, but existing datasets, like TikTok and TED-Talks, lack the diversity and complexity to fully capture the capabilities of video generation models. We close this gap by introducing `What Are You Doing?' (WYD): a new benchmark for fine-grained evaluation of controllable image-to-video generation of humans. WYD consists of 1{,}544 captioned videos that have been meticulously collected and annotated with 56 fine-grained categories. These allow us to systematically measure performance across 9 aspects of human generation, including actions, interactions and motion. We also propose and validate automatic metrics that leverage our annotations and better capture human evaluations. Equipped with our dataset and metrics, we perform in-depth analyses of seven state-of-the-art models in controllable image-to-video generation, showing how WYD provides novel insights about the capabilities of these models. We release our data and code to drive forward progress in human video generation modeling at https://github.com/google-deepmind/wyd-benchmark.

视频生成人类动作评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。