arXiv:2506.13189cs.HCcs.RO2025-06

用语音和手势结合操控机器人,发现语音增效但有延迟,需因人而异。

Gesture First, LLM-Assisted Voice Complement: Exploring Multimodal Robot 'Puppeteer' Teleoperation Via Virtual Counterpart in Augmented Reality

  • 语音负责高层导航,手势做精细操作,分步协作
  • 42人实验表明纯手势更高效,语音+手势有延迟和识别问题
  • 适合有经验用户,需根据技能动态调整多模态策略

通过增强现实(AR)进行机器人远程操控,为更直观的人机交互提供了可能。本文提出一种头戴式AR‘木偶师’系统,用户通过Meta Quest 3与虚拟机器人对手互动,结合大语言模型(LLM)辅助的语音命令与手部手势控制真实机器人。在42名参与者参与的交叉实验中,对比了仅手势(GO)与语音+手势(VG)两种交互方式在基于AR的拾取-放置模式匹配任务中的表现与用户体验(UX)。在VG模式下,语音负责高层导航,手势处理精细操作,呈顺序分工。结果表明,当前时间敏感任务中GO更具可靠性和效率;而VG虽提升灵活性,但引入延迟与识别问题,增加认知负荷。此外分析了先验机器人经验对性能与体验的影响。据此提炼出一套设计指南:将多模态视为需权衡效率、鲁棒性与用户经验的自适应策略,而非默认越多越好。

原文摘要 · Abstract (English)

Robot teleoperation via augmented reality (AR) offers a promising path toward more intuitive human-robot interaction (HRI). We present a head-mounted AR 'puppeteer' system in which users control a physical robot by interacting with its virtual counterpart robot using large language model (LLM)-assisted voice commands and hand-gesture interaction on the Meta Quest 3. In a within-subject user study with 42 participants performing an AR-based robotic pick-and-place pattern-matching task, we empirically compare two interaction conditions: gesture-only (GO) and combined voice+gesture (VG) on performance and user experience (UX). In VG, voice and gesture operate in a sequential role-allocated manner, with voice handling high-level navigation and gesture handling fine manipulation. Our results show that GO currently provides more reliable and efficient control for this time-critical task, while VG introduces additional flexibility but also latency and recognition issues that can increase workload. We additionally analyze how prior robotics expertise differentiates performance and UX across conditions. Based on these findings, we distill a set of design guidelines for AR 'puppeteer' metaphoric robot teleoperation, framing multimodality as an adaptive strategy that must balance efficiency, robustness, and user expertise rather than assuming that additional modalities are universally beneficial.

人机交互语音控制虚拟现实多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。