arXiv:2504.09243cs.RO2025-04被引 7

让机器人实时评估哪种人类协助方式最有效,减少无效干预。

REALM: Real-Time Estimates of Assistance for Learned Models in Human-Robot Interaction

  • 基于机器人行为不确定性,计算不同交互方式的预期信息价值。
  • 实验证明能用最少人类反馈完成任务,尤其在机器人不确定时效果显著。
  • 适合人机协作中希望降低人工负担的研究者与开发者。

多种实时人机交互机制(如遥操作、纠正性输入、离散选择)可提升人机协同效率。然而,现有研究较少探讨如何融合不同方法,也缺乏机器人根据自身对任务的理解主动评估并请求最有效的协助方式。本文提出一种基于机器人策略动作不确定性的协助价值估计方法,通过构建随机策略在交互后微分熵的数学表达式,比较不同交互方式的预期收益。由于各类输入对人力投入的要求不同,我们结合似然惩罚机制,在信息需求与输入成本间实现平衡。仿真与机器人用户实验均表明,该方法可与新兴学习模型(如扩散模型)协同,精准估算协助价值。用户研究表明,该方法能在机器人行为不确定时以最少的人类反馈实现任务完成。

原文摘要 · Abstract (English)

There are a variety of mechanisms (i.e., input types) for real-time human interaction that can facilitate effective human-robot teaming. For example, previous works have shown how teleoperation, corrective, and discrete (i.e., preference over a small number of choices) input can enable robots to complete complex tasks. However, few previous works have looked at combining different methods, and in particular, opportunities for a robot to estimate and elicit the most effective form of assistance given its understanding of a task. In this paper, we propose a method for estimating the value of different human assistance mechanisms based on the action uncertainty of a robot policy. Our key idea is to construct mathematical expressions for the expected post-interaction differential entropy (i.e., uncertainty) of a stochastic robot policy to compare the expected value of different interactions. As each type of human input imposes a different requirement for human involvement, we demonstrate how differential entropy estimates can be combined with a likelihood penalization approach to effectively balance feedback informational needs with the level of required input. We demonstrate evidence of how our approach interfaces with emergent learning models (e.g., a diffusion model) to produce accurate assistance value estimates through both simulation and a robot user study. Our user study results indicate that the proposed approach can enable task completion with minimal human feedback for uncertain robot behaviors.

人机协作不确定性估计实时反馈扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。