用人类反馈强化学习让机器人更自然地做手势。
Generating Natural and Expressive Robot Gestures through Iterative Reinforcement Learning with Human Feedback using LLMs
- 通过人类反馈迭代优化大模型生成手势
- 手势更流畅、有表现力且与对话匹配
- 适合社交机器人交互与人机共融研究
自然且富有表现力的手势对有效的人机沟通至关重要,尤其在仅靠语言难以传达意图时(如指物)。对于人形机器人Pepper而言,生成自然生动的动作对提升人机交互体验和长期接受度尤为关键。然而,现有方法多依赖专家编写的动画,导致动作僵硬,难以适应动态多样环境;而机器学习方法又常难以捕捉自然感,自由度越高越难实现。为此,我们引入大语言模型ChatGPT,使Pepper能根据对话内容实时生成伴随性手势。尽管如此,初始生成的手势仍显呆板。为此,我们提出一种基于人类反馈的迭代强化学习(RLHF)系统,通过多轮用户评估对比,持续优化手势生成策略。实验表明,该系统显著提升了大模型生成手势的表现力、相关性与流畅性。
原文摘要 · Abstract (English)
Expressive gestures are essential for natural and effective communication, complementing speech when verbal cues alone are insufficient (e.g., pointing). For social robots such as the humanoid Pepper, producing natural and expressive movements is critical for improving human-robot interaction (HRI) and long-term acceptance. However, generating gestures remains challenging due to reliance on expert-authored animations, resulting in rigid behaviors that are impractical for dynamic and diverse environments. Alternatively, machine learning approaches often struggle to capture perceived naturalness, becoming increasingly challenging with more degrees of freedom. Consequently, producing expressive robot gestures requires a system that can adapt to the environment while adhering to social norms and physical constraints. Recent advances in large language models (LLMs) enable dynamic code generation, offering new opportunities for runtime gesture synthesis from natural language. In this paper, we integrate ChatGPT into the humanoid robot Pepper to generate co-speech gestures aligned with conversational output. While this baseline enables flexible gesture generation, the resulting motions are often perceived as stiff and unnatural. To address this limitation, we introduce an iterative reinforcement learning with human feedback (RLHF) system that finetunes gesture generation based on user evaluations, leveraging an iterative user study to compare Pepper's generated gestures. Our results show that RLHF improved the LLM's co-speech generative capabilities, producing more expressive, relevant and fluid movements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。