让机器人自评交互体验,发现其对陌生感判断严重失准。
When Robots Rate Their Own Interactions: Engagement Validity and the Strangeness Failure

- 用大模型让机器人从自身视角完成交互问卷
- 机器人对满意度和愉悦度评分与人类一致,但对舒适度判断相反
- 适合研究人机交互评估边界及多模态融合的学者
人机交互评估几乎完全依赖人工问卷,忽略了机器人的视角。本文提出反向评估方法:由大语言模型驱动的机器人以自身视角完成标准化问卷,并检验其评分是否与人类真实评价一致。在研究1中,五个大模型对来自HRI-CUES数据集的25次交互(共1522次评估)完成了HRI-CUES、Godspeed和RoSAS问卷。模型在满意度(相关系数r最高达0.65)和愉悦度(r最高达0.72)维度上达成中到强一致性,且重测信度优异(ICC ≥ 0.82),但在舒适度/陌生感维度上系统性反转(r = -0.44至-0.67,所有p < 0.05),将互动参与度误判为舒适感。研究2中,搭载Claude Sonnet 4.5的Nao机器人在真实互动中(共4名参与者)复现了相同模式,包括实时逐轮评估。该陌生感错误现象在五种模型、合成对照组及两人实体部署中均持续存在。研究认为当前基于大模型的机器人缺乏感知内部情感状态的能力,无法准确评估如陌生感等构念,反向评估需结合生理信号、眼神、空间距离等补充模态,才能突破行为代理的局限。研究明确了大模型作为交互评估者在人机交互中的边界条件。
原文摘要 · Abstract (English)
Human-robot interaction (HRI) evaluation relies almost exclusively on human-completed questionnaires, leaving the robot's perspective unexamined. We propose an \textit{inverted evaluation}, in which LLM-powered robots complete the same standardized instruments from their own perspective, and test whether these ratings agree with human ground truth. In Study~1, five LLMs completed HRI-CUES, Godspeed, and RoSAS questionnaires for 25~interactions ($N = 1{,}522$ evaluations) from the HRI-CUES dataset. LLMs achieved moderate-to-strong agreement on engagement dimensions (satisfaction $r$ up to $.65$ and enjoyment $r$ up to $.72$) with excellent test-retest reliability (ICC $\geq .82$), but \textit{systematically inverted} the comfort/strangeness dimension ($r = -.44$ to $-.67$, all $p < .05$), conflating engagement with comfort. In Study~2, a Nao robot running Claude~Sonnet~4.5 replicated these patterns in live interactions ($N = 4$), including real-time turn-by-turn assessment. The strangeness failure persisted across five models, synthetic controls, and embodied deployment for two participants. We argue that current LLM-based robots lack access to the internal affective states needed to assess constructs like strangeness, and that inverted evaluation requires supplementary modalities (e.g., physiological signals, gaze, proxemics) to move beyond behavioral proxies. These findings establish boundary conditions for using LLMs as interaction evaluators in HRI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。