统一建模语言与动作,实现多轮对话中人互动的生成与理解
A Unified Framework for Motion Reasoning and Generation in Human Interaction
- 构建统一框架,同时处理语言与动作模态的交互生成
- 在82.7K条指令上训练,支持动作编辑、问答与故事生成等任务
- 适用于需要动态响应用户指令的交互系统开发
大语言模型在生成自然文本方面取得显著进展,但生成和理解多人协作的交互式动作仍具挑战。为此,我们提出VIM——一个融合语言与动作模态的通用交互式动作-语言模型,可在多轮对话中理解、生成和控制多人互动动作。不同于以往单向任务(如文到动或动到文),VIM采用统一架构,同步处理双模态。由于缺乏合适数据集,我们构建了包含82.7K条多轮交互动作指令的Inter-MT2数据集,覆盖153K个交互动作样本,涵盖动作编辑、问答与故事生成等多种场景,利用现成的大语言模型与动作扩散模型生成指令。我们在多个交互动作任务(如动到文、文到动、反应生成、动作编辑、动作序列推理)上全面评估VIM的泛化能力。
原文摘要 · Abstract (English)
Recent advancements in large language models (LLMs) have significantly improved their ability to generate natural and contextually relevant text, enabling more human-like AI interactions. However, generating and understanding interactive human-like motion, where multiple individuals engage in coordinated movements, remains challenging due to the complexity of modeling these interactions. Additionally, a unified and versatile model is needed to handle diverse interactive scenarios, such as chat systems that dynamically adapt to user instructions and assigned roles. To address these challenges, we introduce VIM, the Versatile Interactive Motion-language model, which integrates both language and motion modalities to effectively understand, generate, and control interactive motions in multi-turn conversational contexts. Unlike previous studies that primarily focus on uni-directional tasks such as text-to-motion or motion-to-text, VIM employs a unified architecture capable of simultaneously understanding and generating both motion and text modalities. Given the absence of an appropriate dataset to support this task, we introduce Inter-MT2, a large-scale instruction-tuning dataset containing 82.7K multi-turn interactive motion instructions, covering 153K interactive motion samples. Inter-MT2 spans diverse instructional scenarios, including motion editing, question answering, and story generation, leveraging off-the-shelf large language models and motion diffusion models to construct a broad set of interactive motion instructions. We extensively evaluate the versatility of VIM across multiple interactive motion-related tasks, including motion-to-text, text-to-motion, reaction generation, motion editing, and reasoning about motion sequences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。