arXiv:2607.07430cs.ROcs.SY2026-07

用语音和VR操控人形机器人,让新手也能轻松完成远程操作。

Immersive Social Interaction with VR and LLM-Assisted Humanoids

论文配图:Immersive Social Interaction with VR and LLM-Assisted Humanoids
图 1 · 摘自论文原文
  • 通过语音指令和手部追踪实现自然语言控制与精细操作
  • 新手经短时训练即可达80%物品操作成功率、70%社交传递成功率
  • 支持多模态数据采集,适合远程协作与未来自主学习研究

人形机器人可将人类存在延伸至遥远、受限或危险环境,但现有遥控界面常需费力的动作追踪或高认知负荷的底层控制。本文提出一种沉浸式遥控框架,整合语音控制移动、基于虚拟现实的操纵以及双向社交互动,实现全身人形机器人控制。利用Apple Vision Pro,操作者获得第一视角视觉反馈,发出自然语言移动指令,并通过手腕与手指追踪操控机器人手臂及灵巧手。一个由大语言模型辅助的语音控制模块将口语指令转化为高层移动命令,操纵模块则通过逆运动学与PD控制将人类手部动作重定向至机器人。系统还记录多模态数据,包括第一视角RGB观测、语音/文本指令、关节状态、手部动作与眼动信号,支持未来的模仿学习与自主性研究。我们在配备灵巧手的Unitree H1人形机器人上评估该框架,在物品操作与社交传递任务中,新手用户经简短熟悉后分别达到80%与70%的成功率。结果表明,沉浸式语言辅助遥控具有作为人形交互、远程协助与多模态数据采集的可行接口潜力。

原文摘要 · Abstract (English)

Humanoid robots can extend human presence to remote, constrained, or hazardous environments, but existing teleoperation interfaces often require physically demanding motion tracking or cognitively demanding low-level control. This paper presents an immersive teleoperation framework that integrates voice-controlled locomotion, VR-based manipulation, and bidirectional social interaction for whole-body humanoid control. Using Apple Vision Pro, the operator receives egocentric visual feedback, issues natural-language locomotion commands, and teleoperates the robot's arms and dexterous hands through wrist and finger tracking. An LLM-assisted voice-control module converts spoken instructions into high-level locomotion commands, while the manipulation module retargets human hand motions to the robot through inverse kinematics and PD control. The system also records multimodal data, including egocentric RGB observations, voice/text commands, joint states, hand motions, and eye-gaze signals, supporting future imitation learning and autonomy. We evaluate the framework on a Unitree H1 humanoid equipped with dexterous hands in manipulation and social interaction tasks. Results show that novice users can successfully operate the system after brief familiarization, achieving 80\% success in object manipulation and 70\% success in a social cube-passing task. These results demonstrate the potential of immersive, language-assisted teleoperation as an accessible interface for humanoid interaction, remote assistance, and multimodal data collection.

人形机器人语音控制虚拟现实遥操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。