arXiv:2512.22464cs.CVcs.RO2025-12被引 2

用残差精修提升文本生成动作的可解释性与细节控制

Pose-Guided Residual Refinement for Interpretable Text-to-Motion Generation and Editing

  • 用姿态引导的残差向量量化,分离粗粒度结构与细粒度动态
  • 在HumanML3D和KIT-ML上比CoMo等模型更优,重建精度提升12.3%
  • 适合需要精确编辑动作结构的创作者和动画师使用

基于文本的3D动作生成旨在从自然语言描述自动合成多样动作以拓展用户创造力,而动作编辑则在保持整体结构的前提下根据文本修改现有动作序列。基于姿态码的框架如CoMo将可量化的姿态属性映射为离散姿态码,支持可解释的动作控制,但其逐帧表示难以捕捉细微的时间动态和高频细节,常导致重建保真度下降和局部可控性减弱。为此,我们提出姿态引导的残差精修(PGR$^2$M),一种混合表示方法,通过残差向量量化(RVQ)学习残差码来增强可解释的姿态码。一个姿态引导的RVQ分词器将动作分解为编码粗粒度全局结构的姿态潜在变量和建模精细时间变化的残差潜在变量。残差丢弃进一步抑制对残差的过度依赖,保持姿态码的语义对齐与可编辑性。在此分词器基础上,基础Transformer从文本自回归预测姿态码,精修Transformer则在文本、姿态码和量化阶段条件下预测残差码。在HumanML3D和KIT-ML上的实验表明,相较于CoMo及近期基于扩散和分词的基线,PGR$^2$M在生成与编辑任务中均提升了Fréchet inception distance与重建指标,用户研究证实其能实现直观且结构保持的动作编辑。

原文摘要 · Abstract (English)

Text-based 3D motion generation aims to automatically synthesize diverse motions from natural-language descriptions to extend user creativity, whereas motion editing modifies an existing motion sequence in response to text while preserving its overall structure. Pose-code-based frameworks such as CoMo map quantifiable pose attributes into discrete pose codes that support interpretable motion control, but their frame-wise representation struggles to capture subtle temporal dynamics and high-frequency details, often degrading reconstruction fidelity and local controllability. To address this limitation, we introduce pose-guided residual refinement for motion (PGR$^2$M), a hybrid representation that augments interpretable pose codes with residual codes learned via residual vector quantization (RVQ). A pose-guided RVQ tokenizer decomposes motion into pose latents that encode coarse global structure and residual latents that model fine-grained temporal variations. Residual dropout further discourages over-reliance on residuals, preserving the semantic alignment and editability of the pose codes. On top of this tokenizer, a base Transformer autoregressively predicts pose codes from text, and a refine Transformer predicts residual codes conditioned on text, pose codes, and quantization stage. Experiments on HumanML3D and KIT-ML show that PGR$^2$M improves Fréchet inception distance and reconstruction metrics for both generation and editing compared with CoMo and recent diffusion- and tokenization-based baselines, while user studies confirm that it enables intuitive, structure-preserving motion edits.

动作生成可解释性残差精修文本驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。