用大模型实现自然语言控制的沉浸式虚拟现实移动,无需手柄。
Exploring Context-aware and LLM-driven Locomotion for Immersive Virtual Reality
- 基于大语言模型的自然语言指令导航,支持上下文感知。
- 在可用性、沉浸感和晕动症上与传统方式无显著差异。
- 眼动分析显示该方法提升用户注意力与参与度,适合无障碍场景。
运动控制在虚拟现实体验中至关重要。无手柄的运动方式能提升可访问性,摆脱对手持控制器的依赖。传统语音控制常依赖固定命令集,限制了交互的自然性与灵活性。本文提出一种由大语言模型驱动的新型运动技术,使用户可通过自然语言指令实现上下文感知的虚拟空间导航。我们对比了三种运动方式:控制器传送、语音转向,以及本研究提出的语言模型驱动方法。评估结合眼动追踪数据分析(含可解释机器学习的SHAP分析)及标准化问卷(SUS、IPQ、CSQ-VR、NASA-TLX),从客观注视数据和主观报告的可用性、沉浸感、晕动症与认知负荷多维度考察用户体验。结果表明,语言模型驱动方法在可用性、沉浸感与晕动症方面与传送等成熟方法无统计学差异,具备作为自然语言、无手柄替代方案的潜力。此外,眼动分析显示该条件下用户注意力与参与度更高。补充的SHAP分析揭示不同技术下注视、扫视与瞳孔变化模式存在差异,反映视觉注意与认知处理机制的不同。总体而言,该方法可有效支持虚拟空间的无手柄运动,尤其适用于无障碍设计。
原文摘要 · Abstract (English)
Locomotion plays a crucial role in shaping the user experience within virtual reality environments. In particular, hands-free locomotion offers a valuable alternative by supporting accessibility and freeing users from reliance on handheld controllers. To this end, traditional speech-based methods often depend on rigid command sets, limiting the naturalness and flexibility of interaction. In this study, we propose a novel locomotion technique powered by large language models (LLMs), which allows users to navigate virtual environments using natural language with contextual awareness. We evaluate three locomotion methods: controller-based teleportation, voice-based steering, and our language model-driven approach. Our evaluation combines eye-tracking data analysis, including exploratory explainable machine learning analysis with SHAP, and standardized questionnaires (SUS, IPQ, CSQ-VR, NASA-TLX) to examine user experience through both objective gaze-based measures and subjective self-reports of usability, presence, cybersickness, and cognitive load. Our findings show no statistically significant differences in usability, presence, or cybersickness between LLM-driven locomotion and established methods such as teleportation, suggesting its potential as a viable, natural language-based, hands-free alternative. In addition, eye-tracking analysis revealed patterns suggesting tendency toward increased user attention and engagement in the LLM-driven condition. Complementary to these findings, exploratory SHAP analysis revealed that fixation, saccade, and pupil-related features vary across techniques, indicating distinct patterns of visual attention and cognitive processing. Overall, we state that our method can facilitate hands-free locomotion in virtual spaces, especially in supporting accessibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。