arXiv:2511.18927cs.CV2025-11AAAI被引 4

用细粒度文本精准控制身体部位运动,更灵活高效。

FineXtrol: Controllable Motion Generation via Fine-Grained Text

  • 通过分层对比学习增强文本编码器对细粒度指令的区分能力。
  • 在可控动作生成任务中表现优异,能精准引导特定身体部位动作。
  • 适合需要精细动作控制的动画生成、虚拟角色交互等场景。

近期研究致力于提升文本驱动动作生成的可控性与精确度。一些方法利用大语言模型生成更详细的文本描述,另一些则引入全局3D坐标序列作为额外控制信号。然而,前者常导致细节错位且缺乏明确的时间线索,后者在将坐标转换为标准动作表示时带来显著计算开销。为此,我们提出FineXtrol,一种新型控制框架,通过时间感知、精确、用户友好的细粒度文本控制信号,指导动作生成,具体描述随时间变化的身体各部位运动。为支持该框架,我们设计了分层对比学习模块,促使文本编码器对新提出的控制信号生成更具判别性的嵌入表示,从而提升动作可控性。定量结果表明,FineXtrol在可控动作生成任务中表现强劲;定性分析显示其在引导特定身体部位运动方面具有高度灵活性。

原文摘要 · Abstract (English)

Recent works have sought to enhance the controllability and precision of text-driven motion generation. Some approaches leverage large language models (LLMs) to produce more detailed texts, while others incorporate global 3D coordinate sequences as additional control signals. However, the former often introduces misaligned details and lacks explicit temporal cues, and the latter incurs significant computational cost when converting coordinates to standard motion representations. To address these issues, we propose FineXtrol, a novel control framework for efficient motion generation guided by temporally-aware, precise, user-friendly, and fine-grained textual control signals that describe specific body part movements over time. In support of this framework, we design a hierarchical contrastive learning module that encourages the text encoder to produce more discriminative embeddings for our novel control signals, thereby improving motion controllability. Quantitative results show that FineXtrol achieves strong performance in controllable motion generation, while qualitative analysis demonstrates its flexibility in directing specific body part movements.

动作生成文本控制细粒度可控性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。