arXiv:2503.16451cs.HCcs.AI2025-03中稿 · ICLR被引 9

用大模型生成人与人互动的自然反应,提升机器人交互真实感。

Think-Then-React: Towards Unconstrained Human Action-to-Reaction Generation

  • 分两步生成:先推理动作意图,再根据意图生成反应。
  • 在多人体动作生成上,FID降低至1.942,显著优于基线。
  • 适合做人机交互、游戏动画等需要自然反应的场景。

建模类人动作-反应生成在人机交互和游戏等领域有重要应用。尽管单人动作生成已有进展,但直接从动作序列预测反应仍具挑战,主要源于缺乏有效编码多人运动的统一表示。为此,我们提出基于大语言模型的Think-Then-React(TTR)框架。首先,通过细粒度多模态训练策略,TTR在推理时统一两个过程:思考过程显式推断动作意图并生成对应反应描述作为语义提示;反应过程则基于输入动作和推导出的语义提示生成动作。其次,为在语言模型中有效表示多人运动,我们提出一种统一的动作分词器,通过解耦自我中心姿态与绝对空间特征,实现动作与反应的统一编码。大量实验表明,TTR优于现有基线,在评估指标上显著提升,如FID从3.988降至1.942。

原文摘要 · Abstract (English)

Modeling human-like action-to-reaction generation has significant real-world applications, like human-robot interaction and games. Despite recent advancements in single-person motion generation, it is still challenging to well handle action-to-reaction generation, due to the difficulty of directly predicting reaction from action sequence without prompts, and the absence of a unified representation that effectively encodes multi-person motion. To address these challenges, we introduce Think-Then-React (TTR), a large language-model-based framework designed to generate human-like reactions. First, with our fine-grained multimodal training strategy, TTR is capable to unify two processes during inference: a thinking process that explicitly infers action intentions and reasons corresponding reaction description, which serve as semantic prompts, and a reacting process that predicts reactions based on input action and the inferred semantic prompts. Second, to effectively represent multi-person motion in language models, we propose a unified motion tokenizer by decoupling egocentric pose and absolute space features, which effectively represents action and reaction motion with same encoding. Extensive experiments demonstrate that TTR outperforms existing baselines, achieving significant improvements in evaluation metrics, such as reducing FID from 3.988 to 1.942.

动作生成大模型人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。