arXiv:2502.05118cs.ROcs.HC2025-02

研究萌系机器人对人类反馈偏见的影响,提升强化学习中的反馈质量。

Use of Winsome Robots for Understanding Human Feedback (UWU)

  • 通过用户实验分析机器人萌度对反馈倾向的影响。
  • 发现萌度越高,正向反馈比例显著上升,负向反馈减少。
  • 提出自适应算法缓解反馈偏差,适合人机交互与强化学习研究者。

随着社交机器人日益普及,许多机器人采用萌系外观以增强用户舒适感与接受度。然而,这种美学设计对强化学习场景中人类反馈的影响尚不明确。已有研究表明,人类倾向于给出比负面更多的正面反馈,这可能导致机器人无法达到最优行为。我们假设,机器人被感知的萌度可能加剧这一正向偏见。为此,我们开展用户研究,让参与者评估机器人执行任务时的轨迹表现,并分析其外观萌度对反馈类型的影响。结果表明,当感知萌度变化时,正负反馈的比例发生显著转变。基于此,我们尝试改进TAMER算法,引入一种基于用户正向反馈偏见程度的随机化版本,以缓解该问题。

原文摘要 · Abstract (English)

As social robots become more common, many have adopted cute aesthetics aiming to enhance user comfort and acceptance. However, the effect of this aesthetic choice on human feedback in reinforcement learning scenarios remains unclear. Previous research has shown that humans tend to give more positive than negative feedback, which can cause failure to reach optimal robot behavior. We hypothesize that this positive bias may be exacerbated by the robot's level of perceived cuteness. To investigate, we conducted a user study where participants critique a robot's trajectories while it performs a task. We then analyzed the impact of the robot's aesthetic cuteness on the type of participant feedback. Our results suggest that there is a shift in the ratio of positive to negative feedback when perceived cuteness changes. In light of this, we experiment with a stochastic version of TAMER which adapts based on the user's level of positive feedback bias to mitigate these effects.

人机交互强化学习反馈偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。