通过动态混合动作生成训练数据,实现无需LLM的文本控制动作编辑。
Dynamic Motion Blending for Versatile Motion Editing
- 基于输入文本动态拼接身体部位动作,在线生成训练数据。
- 在多个编辑场景下达到当前最优性能,支持空间与时间动作修改。
- 无需额外标注或大语言模型,直接响应高层指令。
文本引导的动作编辑实现了超越传统关键帧动画的高层语义控制和迭代修改。现有方法依赖有限的预收集训练三元组,严重限制了其在多样化编辑场景中的灵活性。我们提出MotionCutMix,一种在线数据增强技术,通过根据输入文本动态混合身体部位动作来生成训练三元组。尽管MotionCutMix有效扩展了训练分布,但组合性质引入了更高的随机性和潜在的身体部位不协调问题。为建模这种丰富分布,我们提出MotionReFit,一种带有动作协调器的自回归扩散模型。自回归架构通过分解长序列促进学习,而动作协调器则缓解动作组合带来的伪影。我们的方法可直接从高层人类指令中处理空间与时间动作编辑,无需额外标注或大语言模型。大量实验表明,MotionReFit在文本引导动作编辑任务上达到当前最佳性能。
原文摘要 · Abstract (English)
Text-guided motion editing enables high-level semantic control and iterative modifications beyond traditional keyframe animation. Existing methods rely on limited pre-collected training triplets, which severely hinders their versatility in diverse editing scenarios. We introduce MotionCutMix, an online data augmentation technique that dynamically generates training triplets by blending body part motions based on input text. While MotionCutMix effectively expands the training distribution, the compositional nature introduces increased randomness and potential body part incoordination. To model such a rich distribution, we present MotionReFit, an auto-regressive diffusion model with a motion coordinator. The auto-regressive architecture facilitates learning by decomposing long sequences, while the motion coordinator mitigates the artifacts of motion composition. Our method handles both spatial and temporal motion edits directly from high-level human instructions, without relying on additional specifications or Large Language Models. Through extensive experiments, we show that MotionReFit achieves state-of-the-art performance in text-guided motion editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。