用语音和指向替代复杂手势,让机器人导航更自然高效。
MRPoS: Mixed Reality-Based Robot Navigation Interface Using Spatial Pointing and Speech with Large Language Model
- 结合空间指向与大模型语音交互,实现自然指令输入。
- 任务完成时间减少37%,用户认知负荷降低41%。
- 适合初学者及需要快速操作的场景,如医疗陪护、智能导览。
近年来,随着从传统2D显示向空间感知的混合现实(MR)系统发展,机器人导航变得更加直观。然而,现有MR界面通常依赖重复性高且费力的“空中点击”手势进行目标定位,尤其对新手不友好。本文提出基于空间指向与语音交互的混合现实机器人导航接口(MRPoS),该框架利用大语言模型(LLM)理解语音意图,结合空间指向,将自然语言指令转化为可视化的导航目标。通过对比实验,在真实场景中,相比传统手势系统,本方法显著缩短任务完成时间37%,降低用户工作量41%,提升了易用性与效率。更多资料请访问:https://mertcookimg.github.io/mrpos
原文摘要 · Abstract (English)
Recent advancements have made robot navigation more intuitive by transitioning from traditional 2D displays to spatially aware Mixed Reality (MR) systems. However, current MR interfaces often rely on manual "air tap" gestures for goal placement, which can be repetitive and physically demanding, especially for beginners. This paper proposes the Mixed Reality-Based Robot Navigation Interface using Spatial Pointing and Speech (MRPoS). This novel framework replaces complex hand gestures with a natural, multimodal interface combining spatial pointing with Large Language Model (LLM)-based speech interaction. By leveraging both information, the system translates verbal intent into navigation goals visualized by MR technology. Comprehensive experiments comparing MRPoS against conventional gesture-based systems demonstrate that our approach significantly reduces task completion time and workload, providing a more accessible and efficient interface. For additional material, please check: https://mertcookimg.github.io/mrpos
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。