无需训练即可生成与场景匹配的文本驱动动作,突破了对海量真实动作数据的依赖。
TSTMotion: Training-free Scene-aware Text-to-motion Generation

- 利用预训练模型推理生成场景感知的动作引导,不需重新训练
- 在多个3D场景上实现自然的动作生成,保持动作连贯性与语义一致性
- 适合缺乏标注数据的研究者快速部署场景感知动作生成系统
文本到动作生成近年来受到广泛关注,主要集中在空白背景下的动作序列生成。然而人类动作通常发生在多样化的三维场景中,这促使研究者探索场景感知的文本到动作生成方法。现有方法多依赖大规模真实动作数据集,而这些数据采集成本高昂。为此,我们首次提出一种无需训练的场景感知文本到动作框架TSTMotion,可高效赋予预训练的空白背景动作生成器场景感知能力。具体地,在给定3D场景和文本描述条件下,结合基础模型推理、预测并验证场景感知动作引导;随后通过两项改进将该引导融入空白背景生成器,得到场景感知的文本驱动动作序列。大量实验表明本框架具有优异性能与泛化能力。代码已开源于项目页面。
原文摘要 · Abstract (English)
Text-to-motion generation has recently garnered significant research interest, primarily focusing on generating human motion sequences in blank backgrounds. However, human motions commonly occur within diverse 3D scenes, which has prompted exploration into scene-aware text-to-motion generation methods. Yet, existing scene-aware methods often rely on large-scale ground-truth motion sequences in diverse 3D scenes, which poses practical challenges due to the expensive cost. To mitigate this challenge, we are the first to propose a \textbf{T}raining-free \textbf{S}cene-aware \textbf{T}ext-to-\textbf{Motion} framework, dubbed as \textbf{TSTMotion}, that efficiently empowers pre-trained blank-background motion generators with the scene-aware capability. Specifically, conditioned on the given 3D scene and text description, we adopt foundation models together to reason, predict and validate a scene-aware motion guidance. Then, the motion guidance is incorporated into the blank-background motion generators with two modifications, resulting in scene-aware text-driven motion sequences. Extensive experiments demonstrate the efficacy and generalizability of our proposed framework. We release our code in \href{https://tstmotion.github.io/}{Project Page}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。