用大模型生成拟人化机器人动作,让互动更自然。
EMOTION: Expressive Motion Sequence Generation for Humanoid Robots with In-Context Learning
- 利用大模型上下文学习能力动态生成手势动作。
- 10种表情动作测试中表现媲美甚至超越真人操作。
- 适合关注人机交互自然性的研究者和开发者。
本文提出EMOTION框架,用于生成类人的拟人化运动序列,提升人形机器人在非语言交流中的表现力。面部表情、手势和身体动作等非语言线索在人际互动中至关重要。尽管机器人行为技术不断进步,现有方法仍难以模仿人类非语言交流的多样性与细微差别。为此,本方法利用大语言模型(LLMs)的上下文学习能力,动态生成适用于人机交互的社会性手势动作序列。我们基于该框架生成了10种不同风格的表达性动作,并通过在线用户研究,将EMOTION及其改进版EMOTION++生成的动作与真人操作的结果进行对比,评估其自然度与可理解性。结果表明,在特定场景下,该方法生成的动作在可理解性和自然度上达到或超过人类水平。同时,研究还为未来设计表达性机器人动作提供了变量参考建议。
原文摘要 · Abstract (English)
This paper introduces a framework, called EMOTION, for generating expressive motion sequences in humanoid robots, enhancing their ability to engage in humanlike non-verbal communication. Non-verbal cues such as facial expressions, gestures, and body movements play a crucial role in effective interpersonal interactions. Despite the advancements in robotic behaviors, existing methods often fall short in mimicking the diversity and subtlety of human non-verbal communication. To address this gap, our approach leverages the in-context learning capability of large language models (LLMs) to dynamically generate socially appropriate gesture motion sequences for human-robot interaction. We use this framework to generate 10 different expressive gestures and conduct online user studies comparing the naturalness and understandability of the motions generated by EMOTION and its human-feedback version, EMOTION++, against those by human operators. The results demonstrate that our approach either matches or surpasses human performance in generating understandable and natural robot motions under certain scenarios. We also provide design implications for future research to consider a set of variables when generating expressive robotic gestures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。