arXiv:2505.21531cs.CVcs.AI2025-05EMNLP被引 2

LLM控制3D虚拟人动作,能理解整体动作但难精准定位肢体。

How Much Do Large Language Models Know about Human Motion? A Case Study in 3D Avatar Control

  • 分两步生成动作:先规划整体动作流程,再细化各部位位置。
  • 对基础动作理解好,但多关节复杂动作定位误差大。
  • 擅长创意动作和文化特有动作,不擅长精确空间描述。

我们通过3D虚拟人控制探索大语言模型(LLMs)对人体运动的认知。给定动作指令后,提示LLMs先生成包含连续步骤的高层运动计划(高阶规划),再在每一步指定身体部位的位置(低阶规划),并通过线性插值生成虚拟人动画。使用20个代表性动作指令,覆盖基本动作并平衡身体部位使用,开展全面评估,包括人工与自动评分高阶计划与生成动画,以及自动对比低阶规划中与理想位置的差异。结果表明,LLMs在理解高层身体动作方面表现强劲,但在精确身体部位定位上存在困难。尽管将动作查询分解为原子成分能提升规划效果,但涉及高自由度身体部位的多步动作仍具挑战。此外,LLMs对一般空间描述给出合理近似,但在处理精确空间规格时表现不足。值得注意的是,LLMs在构思创造性动作和区分文化特异性动作模式方面展现出潜力。

原文摘要 · Abstract (English)

We explore the human motion knowledge of Large Language Models (LLMs) through 3D avatar control. Given a motion instruction, we prompt LLMs to first generate a high-level movement plan with consecutive steps (High-level Planning), then specify body part positions in each step (Low-level Planning), which we linearly interpolate into avatar animations. Using 20 representative motion instructions that cover fundamental movements and balance body part usage, we conduct comprehensive evaluations, including human and automatic scoring of both high-level movement plans and generated animations, as well as automatic comparison with oracle positions in low-level planning. Our findings show that LLMs are strong at interpreting high-level body movements but struggle with precise body part positioning. While decomposing motion queries into atomic components improves planning, LLMs face challenges in multi-step movements involving high-degree-of-freedom body parts. Furthermore, LLMs provide reasonable approximations for general spatial descriptions, but fall short in handling precise spatial specifications. Notably, LLMs demonstrate promise in conceptualizing creative motions and distinguishing culturally specific motion patterns.

3D虚拟人动作生成大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。