arXiv:2508.07540cs.CV2025-08ICCV被引 2

让AI理解抽象指令生成3D人体姿态

CoT-Pose: Chain-of-Thought Reasoning for 3D Pose Generation from Abstract Prompts

  • 引入思维链推理,将抽象语言转化为动作意图
  • 自动生成三元组数据训练模型,提升语义对齐性
  • 适合需要自然语言交互的虚拟角色动画场景

多模态大语言模型与思维链(CoT)推理的进展显著推动了图像和文本生成。然而,3D人体姿态生成仍面临关键挑战:现有文本到姿态模型依赖详细(低层次)提示,明确描述关节位置,而人类沟通常使用抽象(高层次)语言表达动作意图。这种不匹配限制了姿态生成系统在真实场景中的应用。为此,我们提出新框架CoT-Pose,将思维链推理融入姿态生成过程,实现从抽象文本提示生成准确3D人体姿态。我们进一步设计自动数据合成流程,生成包含抽象提示、详细提示和对应3D姿态的三元组数据用于训练。实验表明,该推理增强模型能有效从抽象输入生成合理且语义一致的姿态。本工作强调高层理解在姿态生成中的重要性,并为基于推理的姿势生成开辟新方向。

原文摘要 · Abstract (English)

Recent advances in multi-modal large language models (MLLMs) and chain-of-thought (CoT) reasoning have led to significant progress in image and text generation tasks. However, the field of 3D human pose generation still faces critical limitations. Most existing text-to-pose models rely heavily on detailed (low-level) prompts that explicitly describe joint configurations. In contrast, humans tend to communicate actions and intentions using abstract (high-level) language. This mismatch results in a practical challenge for deploying pose generation systems in real-world scenarios. To bridge this gap, we introduce a novel framework that incorporates CoT reasoning into the pose generation process, enabling the interpretation of abstract prompts into accurate 3D human poses. We further propose a data synthesis pipeline that automatically generates triplets of abstract prompts, detailed prompts, and corresponding 3D poses for training process. Experimental results demonstrate that our reasoning-enhanced model, CoT-Pose, can effectively generate plausible and semantically aligned poses from abstract textual inputs. This work highlights the importance of high-level understanding in pose generation and opens new directions for reasoning-enhanced approach for human pose generation.

3D姿态生成思维链推理自然语言交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。