arXiv:2603.28338cs.HCcs.RO2026-03

对比三种人机交互界面,发现沉浸式远程接口最提升对话自然度和社交感。

Users and Wizards in Conversations: How WoZ Interface Choices Define Human-Robot Interactions

  • 用三种不同限制的巫师界面模拟机器人交互,测试用户与操控者体验差异。
  • 沉浸式虚拟现实界面使用户感知到更强的社交存在感,对话连接最紧密。
  • 该研究建议未来人机交互实验应采用远程沉浸式接口以更真实反映未来机器人场景。

本文研究了巫师-奥兹(WoZ)界面的选择如何影响用户与机器人之间的沟通,从用户和操控者双重视角分析。在对话情境中,使用三种不同程度限制对话输入输出的界面:a)受限感知界面,显示固定视角视频与语音识别转录,操控者只能触发预设话语与动作;b)无限制感知界面,增加实时参与者与机器人的音频;c)VR远程呈现界面,为操控者提供沉浸式立体音视频,并将操控者的即兴言语、注视与面部表情实时传给机器人。结果显示,用户更偏爱由VR界面中介的交互,认为机器人更具社交性且互动更自然。对操控者而言,VR条件最具挑战性但带来更高的情感连接。此外,VR界面产生的说话者间间隙最小、重叠最多,而受限界面则导致最少的互动连贯性与最大的沉默。因此,我们主张更多沃兹研究应采用远程呈现接口,这些接口更能反映未来机器人形态,为基于自然情境下语言与非语言行为数据的自动化提供可行路径。

原文摘要 · Abstract (English)

In this paper, we investigated how the choice of a Wizard-of-Oz (WoZ) interface affects communication with a robot from both the user's and the wizard's perspective. In a conversational setting, we used three WoZ interfaces with varying levels of dialogue input and output restrictions: a) a restricted perception GUI that showed fixed-view video and ASR transcripts and let the wizard trigger pre-scripted utterances and gestures; b) an unrestricted perception GUI that added real-time audio from the participant and the robot c) a VR telepresence interface that streamed immersive stereo video and audio to the wizard and forwarded the wizard's spontaneous speech, gaze and facial expressions to the robot. We found that the interaction mediated by the VR interface was preferred by users in terms of robot features and perceived social presence. For the wizards, the VR condition turned out to be the most demanding but elicited a higher social connection with the users. VR interface also induced the most connected interaction in terms of inter-speaker gaps and overlaps, while Restricted GUI induced the least connected flow and the largest silences. Given these results, we argue for more WoZ studies using telepresence interfaces. These studies better reflect the robots of tomorrow and offer a promising path to automation based on naturalistic contextualized verbal and non-verbal behavioral data.

人机交互虚拟现实对话系统用户体验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。