用大模型辅助语音与指向,提升虚拟现实中多物体选择效率。
Large Language Model-assisted Speech and Pointing Benefits Multiple 3D Object Selection in Virtual Reality
- 结合语音指令与指向交互,利用大模型理解语义意图。
- 24人实验表明,面对多个遮挡目标时,性能优于传统小地图方案。
- 即使物体难描述,也能有效工作,适合复杂沉浸式交互设计。
在虚拟现实中,遮挡物体的选择尤其困难,尤其当涉及多个目标时更为棘手。本文探索利用大语言模型辅助多物体选择任务,提出一种基于多模态语音与射线投射的交互技术(AssistVR)。通过一项包含24名参与者、在不同场景复杂度下的对比用户研究验证该方法。结果表明,在存在多个目标时,AssistVR显著优于以小地图为基础的基线方案。出乎意料的是,即便目标难以口头描述,系统仍表现更优。本研究证明了大语言模型驱动的智能多模态交互系统的可行性与潜力,并为未来沉浸式环境中的智能交互设计提供指导。
原文摘要 · Abstract (English)
Selection of occluded objects is a challenging problem in virtual reality, even more so if multiple objects are involved. With the advent of new artificial intelligence technologies, we explore the possibility of leveraging large language models to assist multi-object selection tasks in virtual reality via a multimodal speech and raycast interaction technique. We validate the findings in a comparative user study (n=24), where participants selected target objects in a virtual reality scene with different levels of scene perplexity. The performance metrics and user experience metrics are compared against a mini-map based occluded object selection technique that serves as the baseline. Results indicate that the introduced technique, AssistVR, outperforms the baseline technique when there are multiple target objects. Contrary to the common belief for speech interfaces, AssistVR was able to outperform the baseline even when the target objects were difficult to reference verbally. This work demonstrates the viability and interaction potential of an intelligent multimodal interactive system powered by large laguage models. Based on the results, we discuss the implications for design of future intelligent multimodal interactive systems in immersive environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。