arXiv:2509.12507cs.ROcs.HC2025-09被引 19

用模仿学习+强化学习生成自然精准的指指点点动作。

Learning to Generate Pointing Gestures in Situated Embodied Conversational Agents

  • 结合模仿与强化学习,从少量动作捕捉数据中训练出自然动作策略。
  • 在虚拟现实中用户测试中,指代准确率和自然度均优于现有监督模型。
  • 适合开发具身对话机器人,尤其关注非语言交互的团队可重点关注。

机器人与智能体研究的核心目标之一是实现物理场景中与人类的自然交流。尽管近期工作多聚焦于语言与语音等口头表达,但非语言沟通对灵活互动至关重要。本文提出一种融合模仿学习与强化学习的框架,用于生成具身代理的指指点点动作。基于小规模动作捕捉数据集,该方法学习到一个运动控制策略,能够生成物理上合理、自然且具有高指代准确性的手势。我们在客观指标及虚拟现实指代游戏中对系统进行了评估,对比了监督学习与检索基线。结果表明,本系统在自然度和准确性方面均优于当前最优的监督模型,凸显了模仿-强化学习在交际手势生成中的潜力,也展示了其在机器人应用中的前景。

原文摘要 · Abstract (English)

One of the main goals of robotics and intelligent agent research is to enable natural communication with humans in physically situated settings. While recent work has focused on verbal modes such as language and speech, non-verbal communication is crucial for flexible interaction. We present a framework for generating pointing gestures in embodied agents by combining imitation and reinforcement learning. Using a small motion capture dataset, our method learns a motor control policy that produces physically valid, naturalistic gestures with high referential accuracy. We evaluate the approach against supervised learning and retrieval baselines in both objective metrics and a virtual reality referential game with human users. Results show that our system achieves higher naturalness and accuracy than state-of-the-art supervised models, highlighting the promise of imitation-RL for communicative gesture generation and its potential application to robots.

具身交互手势生成强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。