arXiv:2605.14269cs.CVcs.AI2026-05被引 1

用物理模拟评估人体动作真实性,让生成视频更自然可信。

PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation

论文配图:PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation
图 1 · 摘自论文原文
  • 将视频中人体重建成3D模型,导入物理引擎逐项评估动作合理性。
  • 在人类评价中比现有方法提升68 Elo分,显著改善动作真实感。
  • 适合关注动作物理合理性的视频生成研究者或开发者。

生成逼真人体动作是视频生成的核心挑战。尽管基于强化学习的后训练提升了视频整体质量,但其在人体动作上的进展受限于无法可靠评估动作真实性的奖励信号。现有视频奖励主要依赖2D感知信号,未显式建模3D身体状态、接触关系和动力学特性,常误判漂浮或违反物理的动作为高质量。为此,我们提出PhyMotion,一种结构化、细粒度的动作奖励机制,将重建的3D人体轨迹置于MuJoCo物理模拟器中,从三个维度评估动作质量:运动学合理性、接触与平衡一致性、动力学可行性。每个维度提供连续可解释的评分信号,精准定位动作中的物理违规点。实验表明,PhyMotion与人类判断的相关性显著优于现有方法。在强化学习后训练中,优化PhyMotion带来更大且更一致的提升,使自回归与双向视频生成器在自动指标和盲评中均实现显著改进(+68 Elo)。消融实验显示三维度互补监督,且仅需小幅训练开销即可保持视频整体质量。

原文摘要 · Abstract (English)

Generating realistic human motion is a central yet unsolved challenge in video generation. While reinforcement learning (RL)-based post-training has driven recent gains in general video quality, extending it to human motion remains bottlenecked by a reward signal that cannot reliably score motion realism. Existing video rewards primarily rely on 2D perceptual signals, without explicitly modeling the 3D body state, contact, and dynamics underlying articulated human motion, and often assign high scores to videos with floating bodies or physically implausible movements. To address this, we propose PhyMotion, a structured, fine-grained motion reward that grounds recovered 3D human trajectories in a physics simulator and evaluates motion quality along multiple dimensions of physical feasibility. Concretely, we recover SMPL body meshes from generated videos, retarget them onto a humanoid in the MuJoCo physics simulator, and evaluate the resulting motion along three axes: kinematic plausibility, contact and balance consistency, and dynamic feasibility. Each component provides a continuous and interpretable signal tied to a specific aspect of motion quality, allowing the reward to capture which aspects of motion are physically correct or violated. Experiments show that PhyMotion achieves stronger correlation with human judgments than existing reward formulations. These gains carry over to RL-based post-training, where optimizing PhyMotion leads to larger and more consistent improvements than optimizing existing rewards, improving motion realism across both autoregressive and bidirectional video generators under both automatic metrics and blind human evaluation (+68 Elo gain). Ablations show that the three axes provide complementary supervision signals, while the reward preserves overall video generation quality with only modest training overhead.

视频生成物理模拟动作评估强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。