用残差向量量化提升动作表示的表达力与解耦性
Making Pose Representations More Expressive and Disentangled via Residual Vector Quantization
- 用残差向量量化融合连续运动特征,增强离散动作码表达
- 在HumanML3D上FID降至0.015,Top-1 R-Precision提升至0.510
- 适合需要精细动作控制与编辑的生成任务
文本到动作生成的进展推动了3D人体动作生成与基于文本的动作控制。可控动作生成(CoMo)依赖动作码表示,但离散动作码难以捕捉细微运动细节,限制表达能力。为此,我们提出一种方法:通过残差向量量化(RVQ)将连续运动特征融入基于动作码的潜在表示中。该设计在保持动作码可解释性和可操控性的基础上,有效捕捉高频等细微运动特征。在HumanML3D数据集上的实验表明,模型将弗雷切特起始距离(FID)从0.041降至0.015,同时将Top-1 R-Precision从0.508提升至0.510。对动作码间方向相似性的定性分析进一步验证了模型在动作编辑中的可控性。
原文摘要 · Abstract (English)
Recent progress in text-to-motion has advanced both 3D human motion generation and text-based motion control. Controllable motion generation (CoMo), which enables intuitive control, typically relies on pose code representations, but discrete pose codes alone cannot capture fine-grained motion details, limiting expressiveness. To overcome this, we propose a method that augments pose code-based latent representations with continuous motion features using residual vector quantization (RVQ). This design preserves the interpretability and manipulability of pose codes while effectively capturing subtle motion characteristics such as high-frequency details. Experiments on the HumanML3D dataset show that our model reduces Frechet inception distance (FID) from 0.041 to 0.015 and improves Top-1 R-Precision from 0.508 to 0.510. Qualitative analysis of pairwise direction similarity between pose codes further confirms the model's controllability for motion editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。