arXiv:2603.02379cs.HCcs.RO2026-03被引 1

机器人通过策略性互动,提升人类合作意愿并改善团队表现。

Strategic Shaping of Human Prosociality: A Latent-State POMDP Framework

  • 将人类合作度建模为随时间演变的隐藏状态,用贝叶斯推断动态感知。
  • 在用户研究中,该策略使团队表现和人类合作行为均优于基线方法。
  • 适合人机协作、社交机器人等需要长期信任建立的场景。

我们提出一种决策理论框架,使机器人在重复交互中战略性地塑造人类被推断出的合作状态。将人类合作性建模为随时间演变的隐状态,机器人通过自身行为(如帮助与信号传递)学习推断并影响该状态,采用带有限观测的隐状态部分可观测马尔可夫决策过程(POMDP)进行形式化,并使用期望最大化算法学习转移与观测动态。由此产生的基于信念的策略在任务目标与社会目标间取得平衡,选择能最大化长期合作结果的动作。我们利用用户研究数据评估该模型,结果显示所学策略在团队绩效和提升人类合作行为方面均优于基线策略。

原文摘要 · Abstract (English)

We propose a decision-theoretic framework in which a robot strategically can shape inferred human's prosocial state during repeated interactions. Modeling the human's prosociality as a latent state that evolves over time, the robot learns to infer and influence this state through its own actions, including helping and signaling. We formalize this as a latent-state POMDP with limited observations and learn the transition and observation dynamics using expectation maximization. The resulting belief-based policy balances task and social objectives, selecting actions that maximize long-term cooperative outcomes. We evaluate the model using data from user studies and show that the learned policy outperforms baseline strategies in both team performance and increasing observed human cooperative behavior.

人机协作强化学习合作行为贝叶斯推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。