arXiv:2509.16858cs.RO2025-09

用离线强化学习让机器人在游戏互动中自适应感知情绪并调整行为。

Towards an Adaptive Social Game-Playing Robot: An Offline Reinforcement Learning-Based Framework

  • 基于多模态情绪识别与离线强化学习构建自适应交互框架
  • BCQ和DDQN对超参数变化最鲁棒,CQL有效缓解值函数过高估计
  • 适合需要安全、高效训练的社会机器人研发者参考

人机交互研究日益要求机器人超越任务执行,能有意义地回应用户情绪,尤其在支持有学习困难学生的游戏化学习场景中。此时机器人的目标是培养用户的游戏技巧,需获取其兴趣与参与度反馈。本文提出一种自适应社交游戏机器人系统。由于在线强化学习需大量真实世界数据且可能令用户不适,我们探索离线强化学习作为更安全高效的替代方案。系统整合多模态情绪识别与自适应机器人响应机制,并在真实人机游戏交互数据集上评估多种离线强化学习算法性能。结果表明,BCQ与DDQN对超参数变化最具鲁棒性,而CQL在缓解值函数过高估计方面最有效。本研究旨在为实际社会机器人设计可靠离线强化学习策略提供依据,推动完全基于离线数据学习复杂情感自适应行为的智能体发展,兼顾人类舒适性与可扩展性。

原文摘要 · Abstract (English)

HRI research increasingly demands robots that go beyond task execution to respond meaningfully to user emotions. This is especially needed when supporting students with learning difficulties in game-based learning scenarios. Here, the objective of these robots is to train users with game-playing skills, and this requires robots to get input about users' interests and engagement. In this paper, we present a system for an adaptive social game-playing robot. However, creating such an agent through online RL requires extensive real-world training data and potentially be uncomfortable for users. To address this, we investigate offline RL as a safe and efficient alternative. We introduce a system architecture that integrates multimodal emotion recognition and adaptive robotic responses. We also evaluate the performance of various offline RL algorithms using a dataset collected from a real-world human-robot game-playing scenario. Our results indicate that BCQ and DDQN offer the greatest robustness to hyperparameter variations, whereas CQL is the most effective at mitigating overestimation bias. Through this research, we aim to inform the selection and design of reliable offline RL policies for real-world social robotics. Ultimately, this work provides a foundational step toward creating socially intelligent agents that can learn complex and emotion-adaptive behaviors entirely from offline datasets, ensuring both human comfort and practical scalability.

社会机器人离线强化学习情绪识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。