arXiv:2601.01969cs.RO2026-01

通过真实场景实验,对比不同奖励机制对机器人对话策略的影响。

What you reward is what you learn: Comparing rewards for online speech policy optimization in public HRI

  • 将对话策略优化建模为多臂赌博机问题,用汤普森采样选择语速与简洁度组合。
  • 三种奖励(用户评分、对话结束率、互动轮次)导致不同行为分布,效果各异。
  • 适用于希望在真实公共场景中优化人机对话的机器人研发者。

在开放多变的环境中设计高效且可接受的对话服务机器人策略极具挑战。相比固定的手动参数,在线学习能适应非平稳条件。本文研究如何在真实场景中优化社交机器人的对话策略。在为期12天的现场部署中,机器人经历了超过1,400次公众互动,将在线策略优化建模为多臂赌博机问题,并使用汤普森采样在六种策略(语速:慢/正常/快;冗余度:简洁/详细)间进行选择。比较了三种互补的二值奖励:用户评分(Ru)、对话关闭率(Rc)和持续至少两轮对话(Rt)。结果显示,每种奖励均引致不同的动作分布与交互行为。结合视频标注数据的离线分析进一步探讨了人群密度、群体规模等上下文因素的影响。综合结果提炼出可直接用于真实公共人机交互场景中在线优化对话策略的设计建议。

原文摘要 · Abstract (English)

Designing policies that are both efficient and acceptable for conversational service robots in open and diverse environments is non-trivial. Unlike fixed, hand-tuned parameters, online learning can adapt to non-stationary conditions. In this paper, we study how to adapt a social robot's speech policy in the wild. During a 12-day in-situ deployment with over 1,400 public encounters, we cast online policy optimization as a multi-armed bandit problem and use Thompson sampling to select among six actions defined by speech rate (slow/normal/fast) and verbosity (concise/detailed). We compare three complementary binary rewards--Ru (user rating), Rc (conversation closure), and Rt (>=2 turns)--and show that each induces distinct arm distributions and interaction behaviors. We complement the online results with offline evaluations that analyze contextual factors (e.g., crowd level, group size) using video-annotated data. Taken together, we distill ready-to-use design lessons for deploying online optimization of speech policies in real public HRI settings.

人机交互在线学习对话策略强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。