arXiv:2603.09611cs.CV2026-03中稿 · CVPR

让文字生成动作更精准,身体各部位表达更自然。

ParTY: Part-Guidance for Expressive Text-to-Motion Synthesis

  • 先生成局部动作再融合,提升身体部位表达精度。
  • 文本嵌入动态变换并精准对齐身体部位,增强语义匹配。
  • 适合需要精细肢体控制的动作生成场景,如动画、游戏。

文本到动作合成旨在从文本描述生成自然且富有表现力的人体动作。现有方法主要关注整体动作生成,难以准确体现涉及特定身体部位的动作。近期的分部位动作生成方法虽有所改进,但仍存在两大问题:(i) 缺乏显式机制将文本语义与个体身体部位对齐;(ii) 因独立生成各部位动作导致全身动作不连贯。为克服这些局限并解决现有方法的根本矛盾,本文提出 ParTY 框架,在保持全身动作连贯性的同时提升局部表达能力。ParTY 包含:(1) 部分引导网络,先生成局部动作以获取引导信号,再用于生成整体动作;(2) 部分感知文本定位,对文本嵌入进行多样化变换,并合理对齐至各身体部位;(3) 整体-局部融合模块,自适应融合整体动作与局部动作。大量实验,包括局部精度和连贯性评估,证明 ParTY 在多个指标上显著优于先前方法。

原文摘要 · Abstract (English)

Text-to-motion synthesis aims to generate natural and expressive human motions from textual descriptions. While existing approaches primarily focus on generating holistic motions from text descriptions, they struggle to accurately reflect actions involving specific body parts. Recent part-wise motion generation methods attempt to resolve this but face two critical limitations: (i) they lack explicit mechanisms for aligning textual semantics with individual body parts, and (ii) they often generate incoherent full-body motions due to integrating independently generated part motions. To overcome these issues and resolve the fundamental trade-off in existing methods, we propose ParTY, a novel framework that enhances part expressiveness while generating coherent full-body motions. ParTY comprises: (1) Part-Guided Network, which first generates part motions to obtain part guidance, then uses it to generate holistic motions; (2) Part-aware Text Grounding, which diversely transforms text embeddings and appropriately aligns them with each body part; and (3) Holistic-Part Fusion, which adaptively fuses holistic motions and part motions. Extensive experiments, including part-level and coherence-level evaluations, demonstrate that ParTY achieves substantial improvements over previous methods.

动作生成文本驱动身体部位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。