arXiv:2601.03517cs.CV2026-01被引 1

用隐状态模拟人体运动动力学,实现更稳定长时预测。

Semantic Belief-State World Model for 3D Human Motion Prediction

  • 将运动预测转为在人体流形上的隐状态动态演化,分离动作与外观重建。
  • 长时滚动预测保持连贯性,计算开销比现有方法低得多。
  • 适合需要高稳定性运动模拟的机器人、动画生成场景。

人类运动预测传统上被视为序列回归问题,模型从历史姿态中外推未来关节坐标。尽管短时预测有效,但该方法未区分观测重建与动力学建模,且缺乏对运动驱动因素的显式表征,导致超出训练范围后出现累积漂移、均值姿态坍塌和不确定性校准不足。本文提出语义信念状态世界模型(SBWM),将运动预测重构为对人体身体流形的隐状态动态模拟。不直接预测姿态,而是维护一个递归概率信念状态,其演化独立于姿态重建,并显式对齐SMPL-X解剖参数化。这种对齐施加结构信息瓶颈,防止隐状态编码静态几何或传感器噪声,迫使它捕捉运动动力学、意图与控制相关结构。受基于模型强化学习中的信念状态世界模型启发,SBWM采用随机隐状态转移与以滚动为中心的训练策略。相比侧重重建保真度的RSSM、Transformer和扩散模型,SBWM优先保证前向模拟稳定性。实验表明其具备连贯的长时滚动预测能力,在计算成本显著更低的情况下仍保持竞争力。结果表明,将人体视为世界模型状态空间而非输出,从根本上改变了运动的模拟与预测方式。

原文摘要 · Abstract (English)

Human motion prediction has traditionally been framed as a sequence regression problem where models extrapolate future joint coordinates from observed pose histories. While effective over short horizons this approach does not separate observation reconstruction with dynamics modeling and offers no explicit representation of the latent causes governing motion. As a result, existing methods exhibit compounding drift, mean-pose collapse, and poorly calibrated uncertainty when rolled forward beyond the training regime. Here we propose a Semantic Belief-State World Model (SBWM) that reframes human motion prediction as latent dynamical simulation on the human body manifold. Rather than predicting poses directly, SBWM maintains a recurrent probabilistic belief state whose evolution is learned independently of pose reconstruction and explicitly aligned with the SMPL-X anatomical parameterization. This alignment imposes a structural information bottleneck that prevents the latent state from encoding static geometry or sensor noise, forcing it to capture motion dynamics, intent, and control-relevant structure. Inspired by belief-state world models developed for model-based reinforcement learning, SBWM adapts stochastic latent transitions and rollout-centric training to the domain of human motion. In contrast to RSSM-based, transformer, and diffusion approaches optimized for reconstruction fidelity, SBWM prioritizes stable forward simulation. We demonstrate coherent long-horizon rollouts, and competitive accuracy at substantially lower computational cost. These results suggest that treating the human body as part of the world models state space rather than its output fundamentally changes how motion is simulated, and predicted.

运动预测世界模型隐状态长时模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。