用文字描述生成逼真的人类反应动作,让角色互动更自然。
MoReact: Generating Reactive Motion from Textual Descriptions
- 分步生成全局轨迹和局部动作,提升动作连贯性
- 结合文本描述生成多样且符合情境的反应动作
- 适合虚拟角色动画、人机交互等需要动态响应的场景
建模与生成人类反应是计算机视觉与人机交互中的关键挑战。现有方法要么将多人视为单一实体直接生成交互,要么仅依赖一人动作生成另一人反应,未能充分融合互动背后的丰富语义信息。此类方法在自适应响应能力上表现不足,难以应对多样化动态场景。为此,本文提出MoReact,一种基于扩散模型的方法,通过分步生成全局运动轨迹与局部动作,实现对文本描述的精准响应。该方法首先生成全局轨迹以引导局部动作,确保与对手动作及文本描述的一致性。同时引入新颖的交互损失函数,增强近距离交互的真实感。实验基于双人动作数据集进行,结果表明该方法能生成逼真、多样且可控的反应动作,不仅高度匹配对手动作,还严格遵循文本指导。
原文摘要 · Abstract (English)
Modeling and generating human reactions poses a significant challenge with broad applications for computer vision and human-computer interaction. Existing methods either treat multiple individuals as a single entity, directly generating interactions, or rely solely on one person's motion to generate the other's reaction, failing to integrate the rich semantic information that underpins human interactions. Yet, these methods often fall short in adaptive responsiveness, i.e., the ability to accurately respond to diverse and dynamic interaction scenarios. Recognizing this gap, our work introduces an approach tailored to address the limitations of existing models by focusing on text-driven human reaction generation. Our model specifically generates realistic motion sequences for individuals that responding to the other's actions based on a descriptive text of the interaction scenario. The goal is to produce motion sequences that not only complement the opponent's movements but also semantically fit the described interactions. To achieve this, we present MoReact, a diffusion-based method designed to disentangle the generation of global trajectories and local motions sequentially. This approach stems from the observation that generating global trajectories first is crucial for guiding local motion, ensuring better alignment with given action and text. Furthermore, we introduce a novel interaction loss to enhance the realism of generated close interactions. Our experiments, utilizing data adapted from a two-person motion dataset, demonstrate the efficacy of our approach for this novel task, which is capable of producing realistic, diverse, and controllable reactions that not only closely match the movements of the counterpart but also adhere to the textual guidance. Please find our webpage at https://xiyan-xu.github.io/MoReactWebPage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。