测试大模型能否用语言指令操控虚拟现实游戏设备
ComboBench: Can LLMs Manipulate Physical Devices to Play Virtual Reality Games?
- 设计262个场景的基准测试,评估大模型将语义指令转为操作序列的能力
- 顶级模型如Gemini-1.5-Pro仍落后于人类,尤其在空间推理上
- 少量示例能显著提升表现,适合研究人机交互与具身智能的学者
虚拟现实(VR)游戏要求玩家将高层语义动作转化为控制器和头戴设备的精确操作。尽管人类可基于常识和身体感知直观完成这一转换,但大型语言模型(LLMs)是否具备类似能力尚未被充分探索。本文提出一个名为ComboBench的基准,评估7个LLM(包括GPT-3.5、GPT-4、GPT-4o、Gemini-1.5-Pro、LLaMA-3-8B、Mixtral-8x7B、GLM-4-Flash)在四款流行VR游戏——Half-Life: Alyx、Into the Radius、Moss: Book II和Vivecraft中的表现,涵盖262个场景。模型性能与人工标注真值及人类表现对比。结果显示,虽顶级模型如Gemini-1.5-Pro展现出较强的任务分解能力,但在过程推理和空间理解方面仍显著弱于人类。不同游戏间表现差异明显,表明对交互复杂度敏感。少量示例显著提升性能,提示可通过针对性训练增强模型在VR操作中的能力。所有材料已公开于https://sites.google.com/view/combobench。
原文摘要 · Abstract (English)
Virtual Reality (VR) games require players to translate high-level semantic actions into precise device manipulations using controllers and head-mounted displays (HMDs). While humans intuitively perform this translation based on common sense and embodied understanding, whether Large Language Models (LLMs) can effectively replicate this ability remains underexplored. This paper introduces a benchmark, ComboBench, evaluating LLMs' capability to translate semantic actions into VR device manipulation sequences across 262 scenarios from four popular VR games: Half-Life: Alyx, Into the Radius, Moss: Book II, and Vivecraft. We evaluate seven LLMs, including GPT-3.5, GPT-4, GPT-4o, Gemini-1.5-Pro, LLaMA-3-8B, Mixtral-8x7B, and GLM-4-Flash, compared against annotated ground truth and human performance. Our results reveal that while top-performing models like Gemini-1.5-Pro demonstrate strong task decomposition capabilities, they still struggle with procedural reasoning and spatial understanding compared to humans. Performance varies significantly across games, suggesting sensitivity to interaction complexity. Few-shot examples substantially improve performance, indicating potential for targeted enhancement of LLMs' VR manipulation capabilities. We release all materials at https://sites.google.com/view/combobench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。