用黎曼几何统一建模人体动作,生成更自然流畅。
Riemannian Motion Generation: A Unified Framework for Human Motion Representation and Generation via Riemannian Flow Matching
- 在黎曼流形上表示动作,通过测地线插值训练
- 在HumanML3D上FID达0.043,优于现有方法
- 适合需要高保真动作生成的科研与应用
人体动作生成通常在欧几里得空间中学习,但有效动作遵循非欧几何结构。我们提出黎曼动作生成(RMG),一个统一框架,将动作表示在乘积流形上,并通过黎曼流匹配学习动态。RMG将动作分解为多个流形因子,实现无量纲表示与内在归一化;训练和采样时采用测地线插值、切空间监督及流形保持的常微分方程积分。在HumanML3D数据集上,RMG在HumanML3D格式下取得最优FID(0.043),并在MotionStreamer格式下所有指标排名第一。在MotionMillion上,亦超越强基线(FID 5.6,R@1 0.86)。消融实验表明,紧凑的$/mathscr{T}+/mathscr{R}$(平移+旋转)表示最稳定有效,凸显几何感知建模是实现高保真动作生成的可行且可扩展路径。
原文摘要 · Abstract (English)
Human motion generation is often learned in Euclidean spaces, although valid motions follow structured non-Euclidean geometry. We present Riemannian Motion Generation (RMG), a unified framework that represents motion on a product manifold and learns dynamics via Riemannian flow matching. RMG factorizes motion into several manifold factors, yielding a scale-free representation with intrinsic normalization, and uses geodesic interpolation, tangent-space supervision, and manifold-preserving ODE integration for training and sampling. On HumanML3D, RMG achieves state-of-the-art FID in the HumanML3D format (0.043) and ranks first on all reported metrics under the MotionStreamer format. On MotionMillion, it also surpasses strong baselines (FID 5.6, R@1 0.86). Ablations show that the compact $\mathscr{T}+\mathscr{R}$ (translation + rotations) representation is the most stable and effective, highlighting geometry-aware modeling as a practical and scalable route to high-fidelity motion generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。