arXiv:2601.14037cs.CV2026-01被引 1

用现成检测模型做奖励函数,显著提升视频中人体动作生成质量

Human detectors are surprisingly powerful reward models

  • 用人体检测置信度和时序提示对齐分数构造简单奖励函数
  • 无需训练,在复杂人体动作上胜过微调的专用模型,胜率73%
  • 不仅提升人形动作,还改善动物和人物交互视频生成

视频生成模型虽在视觉保真度和时间连贯性上表现优异,但在合成动态人体动作(如运动、舞蹈)等非刚性运动时仍面临挑战,常出现肢体缺失、多余或姿态扭曲等物理不合理现象。本文提出一种简单有效的奖励模型HuDA,通过融合人体检测置信度评估外观质量,结合时序提示对齐分数捕捉动作真实性。该方法仅使用现成模型,无需额外训练,却在后处理阶段的群体奖励策略优化(GRPO)中显著提升生成效果,尤其在复杂人体动作上超越主流模型Wan 2.1,胜率达到73%。进一步实验表明,该奖励机制还可泛化至动物视频与人-物交互场景,有效提升整体生成质量。

原文摘要 · Abstract (English)

Video generation models have recently achieved impressive visual fidelity and temporal coherence. Yet, they continue to struggle with complex, non-rigid motions, especially when synthesizing humans performing dynamic actions such as sports, dance, etc. Generated videos often exhibit missing or extra limbs, distorted poses, or physically implausible actions. In this work, we propose a remarkably simple reward model, HuDA, to quantify and improve the human motion in generated videos. HuDA integrates human detection confidence for appearance quality, and a temporal prompt alignment score to capture motion realism. We show this simple reward function that leverages off-the-shelf models without any additional training, outperforms specialized models finetuned with manually annotated data. Using HuDA for Group Reward Policy Optimization (GRPO) post-training of video models, we significantly enhance video generation, especially when generating complex human motions, outperforming state-of-the-art models like Wan 2.1, with win-rate of 73%. Finally, we demonstrate that HuDA improves generation quality beyond just humans, for instance, significantly improving generation of animal videos and human-object interactions.

视频生成人体动作奖励模型扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。