仅凭一句文字指令和目标位置,自动生成角色在场景中的多阶段动作。
Autonomous Character-Scene Interaction Synthesis from Text Instruction
- 用自回归扩散模型逐段生成动作,结合智能调度器自动切换动作阶段。
- 在120个室内场景中生成40类动作,与文本和环境高度匹配。
- 适合动画制作、游戏开发等需要自动化角色行为的领域。
在3D环境中合成人类动作,特别是涉及行走、伸手及人物-物体交互等复杂活动时,现有方法依赖大量手动设定的关键点和阶段转换,难以实现从简单输入自动完成角色动画。本文提出一个端到端框架,仅需一句文本指令和目标位置,即可生成多阶段、场景感知的动作序列。采用自回归扩散模型生成下一动作片段,并引入自主调度器预测各动作阶段的转换。为确保动作与环境融合自然,设计了同时考虑起始与目标位置局部感知的场景表示方法。通过将帧嵌入与语言输入融合,提升生成动作的连贯性。此外,构建了一个包含16小时动作数据的综合动捕数据集,覆盖120个室内场景和40种动作类型,每段动作均有精确的语言标注。实验表明,该方法能生成高质量、符合文本与环境条件的多阶段动作。
原文摘要 · Abstract (English)
Synthesizing human motions in 3D environments, particularly those with complex activities such as locomotion, hand-reaching, and human-object interaction, presents substantial demands for user-defined waypoints and stage transitions. These requirements pose challenges for current models, leading to a notable gap in automating the animation of characters from simple human inputs. This paper addresses this challenge by introducing a comprehensive framework for synthesizing multi-stage scene-aware interaction motions directly from a single text instruction and goal location. Our approach employs an auto-regressive diffusion model to synthesize the next motion segment, along with an autonomous scheduler predicting the transition for each action stage. To ensure that the synthesized motions are seamlessly integrated within the environment, we propose a scene representation that considers the local perception both at the start and the goal location. We further enhance the coherence of the generated motion by integrating frame embeddings with language input. Additionally, to support model training, we present a comprehensive motion-captured dataset comprising 16 hours of motion sequences in 120 indoor scenes covering 40 types of motions, each annotated with precise language descriptions. Experimental results demonstrate the efficacy of our method in generating high-quality, multi-stage motions closely aligned with environmental and textual conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。