arXiv:2411.01866cs.ROcs.AI2024-11被引 2

用连续奖励优化信任估计,让机器人实时感知人类信任变化。

Improving Trust Estimation in Human-Robot Collaboration Using Beta Reputation at Fine-grained Timescales

  • 用每步更新的连续奖励替代任务后才更新的二元反馈
  • 通过最大熵优化自动构建奖励函数,精度提升明显
  • 适合需要实时适应的人机协作场景,如手术辅助

人机协作中,人类会根据感知到的信任调整行为。为实现类似适应性,机器人需在细粒度时间尺度上准确估计人类信任。传统贝塔信誉模型依赖任务级二元结果,仅在任务结束后更新信任,且需人工设计奖励函数,耗时费力。本文提出一种基于贝塔信誉的细粒度信任估计框架,通过每时刻使用连续奖励值更新信任,并采用最大熵优化自动构造奖励函数,避免人工干预。该方法显著提升信任估计准确性,无需手动设计性能指标,推动更智能人机协作机器人的发展。

原文摘要 · Abstract (English)

When interacting with each other, humans adjust their behavior based on perceived trust. To achieve similar adaptability, robots must accurately estimate human trust at sufficiently granular timescales while collaborating with humans. Beta reputation is a popular way to formalize a mathematical estimation of human trust. However, it relies on binary performance, which updates trust estimations only after each task concludes. Additionally, manually crafting a reward function is the usual method of building a performance indicator, which is labor-intensive and time-consuming. These limitations prevent efficient capture of continuous trust changes at more granular timescales throughout the collaboration task. Therefore, this paper presents a new framework for the estimation of human trust using beta reputation at fine-grained timescales. To achieve granularity in beta reputation, we utilize continuous reward values to update trust estimates at each timestep of a task. We construct a continuous reward function using maximum entropy optimization to eliminate the need for the laborious specification of a performance indicator. The proposed framework improves trust estimations by increasing accuracy, eliminating the need to manually craft a reward function, and advancing toward the development of more intelligent robots.

人机协作信任估计贝塔信誉强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。