用拉班动作分析引导扩散模型生成有表现力的多样化动作。
LaMoGen: Laban Movement-Guided Diffusion for Text-to-Motion Generation
- 将拉班动作的力度与形态量化融入文本驱动动作生成。
- 无需额外数据,仅通过优化文本嵌入实现动作属性控制。
- 适合需要精细动作控制的舞蹈生成与动画设计场景。
多样化的动作生成在计算机视觉、人机交互和动画领域具有重要意义。尽管基于扩散模型的文本到动作合成已能生成高质量动作,但实现细粒度的表现力控制仍面临挑战,主要源于数据集中动作风格多样性不足以及自然语言难以表达量化特征。拉班动作分析被舞者广泛用于描述动作细节,包括动作质量的一致性表达。受此启发,本文提出一种可解释且具表现力的动作生成控制方法,通过无缝整合拉班动作的力度(Effort)与形态(Shape)量化方法至文本引导的动作生成模型中。所提方法为零样本、推理时优化策略,在采样过程中仅通过更新预训练扩散模型的文本嵌入,即可实现目标拉班标签对应的动作属性控制,无需额外动作数据。实验表明,该方法在保持动作身份一致性的前提下,成功生成多样化且富有表现力的动作质量。
原文摘要 · Abstract (English)
Diverse human motion generation is an increasingly important task, having various applications in computer vision, human-computer interaction and animation. While text-to-motion synthesis using diffusion models has shown success in generating high-quality motions, achieving fine-grained expressive motion control remains a significant challenge. This is due to the lack of motion style diversity in datasets and the difficulty of expressing quantitative characteristics in natural language. Laban movement analysis has been widely used by dance experts to express the details of motion including motion quality as consistent as possible. Inspired by that, this work aims for interpretable and expressive control of human motion generation by seamlessly integrating the quantification methods of Laban Effort and Shape components into the text-guided motion generation models. Our proposed zero-shot, inference-time optimization method guides the motion generation model to have desired Laban Effort and Shape components without any additional motion data by updating the text embedding of pretrained diffusion models during the sampling step. We demonstrate that our approach yields diverse expressive motion qualities while preserving motion identity by successfully manipulating motion attributes according to target Laban tags.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。