arXiv:2603.11605cs.CV2026-03中稿 · CVPR被引 1

用符号推理让大模型生成可解释、精准的肢体动作。

LaMoGen: Language to Motion Generation Through LLM-Guided Symbolic Inference

  • 用改良版拉班记谱法拆解动作,实现语言到动作的符号映射。
  • 在三个维度上超越现有方法,动作与描述匹配度显著提升。
  • 适合需要可控、可解释动作生成的研究者和创作者。

人类动作高度具表现力且自然契合语言,但现有依赖文本-动作嵌入的方法难以生成时间精确、细节丰富的动作,且缺乏可解释性。为此,我们提出 LabanLite,一种基于拉班记谱法改进的动作表示体系,将每个基本身体动作(如单次左脚迈步)编码为离散的拉班符号与文本模板的组合。该抽象将复杂动作分解为可解释的符号序列与部位指令,建立高层语言与底层运动轨迹之间的符号关联。基于 LabanLite,我们构建 LaMoGen 框架,使大语言模型通过符号推理生成动作序列:理解动作模式,关联文本描述,并重组符号形成可执行计划,生成兼具可解释性与语言对齐性的动作。为支持严谨评估,我们引入基于拉班记谱法的基准数据集,包含结构化描述-动作对及三项指标,综合衡量符号、时间与和谐维度上的文本-动作对齐。实验表明,LaMoGen 在可解释性与可控性上建立新基准,在本研究基准及两个公开数据集上均优于现有方法。结果凸显了符号推理与基于代理设计在语言驱动动作生成中的优势。

原文摘要 · Abstract (English)

Human motion is highly expressive and naturally aligned with language, yet prevailing methods relying heavily on joint text-motion embeddings struggle to synthesize temporally accurate, detailed motions and often lack explainability. To address these limitations, we introduce LabanLite, a motion representation developed by adapting and extending the Labanotation system. Unlike black-box text-motion embeddings, LabanLite encodes each atomic body-part action (e.g., a single left-foot step) as a discrete Laban symbol paired with a textual template. This abstraction decomposes complex motions into interpretable symbol sequences and body-part instructions, establishing a symbolic link between high-level language and low-level motion trajectories. Building on LabanLite, we present LaMoGen, a Text-to-LabanLite-to-Motion Generation framework that enables large language models (LLMs) to compose motion sequences through symbolic reasoning. The LLM interprets motion patterns, relates them to textual descriptions, and recombines symbols into executable plans, producing motions that are both interpretable and linguistically grounded. To support rigorous evaluation, we introduce a Labanotation-based benchmark with structured description-motion pairs and three metrics that jointly measure text-motion alignment across symbolic, temporal, and harmony dimensions. Experiments demonstrate that LaMoGen establishes a new baseline for both interpretability and controllability, outperforming prior methods on our benchmark and two public datasets. These results highlight the advantages of symbolic reasoning and agent-based design for language-driven motion synthesis.

动作生成符号推理大模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。