arXiv:2601.07553cs.AI2026-01AAAI

构建沉浸式交互环境,评估大模型在真实场景中的智能表现。

VirtualEnv: A Platform for Embodied AI Research

论文配图:VirtualEnv: A Platform for Embodied AI Research
图 1 · 摘自论文原文
  • 基于虚幻引擎5打造可交互仿真平台,支持物体操作与多智能体协作。
  • 通过自然语言指令控制智能体,实现复杂任务的动态生成与实时验证。
  • 适合研究大模型推理、规划与多智能体协同的学者及游戏AI开发者。

随着大语言模型(LLMs)在推理与决策能力上的持续提升,亟需真实且互动的环境来严格评估其性能。我们提出 VirtualEnv,一个基于 Unreal Engine 5 的下一代仿真平台,支持对 LLM 在具身与交互场景中的细粒度基准测试。该平台具备丰富的智能体-环境交互能力,包括物体操作、导航、自适应多智能体协作,以及密室逃脱等游戏化机制和程序生成环境。我们提供基于虚幻引擎的易用 API,使研究人员可通过自然语言指令部署和控制由 LLM 驱动的智能体。平台集成 GPT 系列等大规模 LLM 与视觉-语言模型(VLMs),能从多模态输入生成新颖环境与结构化任务。实验在逐步增加复杂性的任务中评测多个主流 LLM,分析其适应性、规划能力与多智能体协调差异。我们还介绍了程序化任务生成、任务验证与实时环境控制的方法。VirtualEnv 已开源,旨在推动人工智能与游戏交叉研究,实现具身 AI 场景下 LLM 评估的标准化,并为沉浸式仿真与交互娱乐的未来发展铺路。

原文摘要 · Abstract (English)

As large language models (LLMs) continue to improve in reasoning and decision-making, there is a growing need for realistic and interactive environments where their abilities can be rigorously evaluated. We present VirtualEnv, a next-generation simulation platform built on Unreal Engine 5 that enables fine-grained benchmarking of LLMs in embodied and interactive scenarios. VirtualEnv supports rich agent-environment interactions, including object manipulation, navigation, and adaptive multi-agent collaboration, as well as game-inspired mechanics like escape rooms and procedurally generated environments. We provide a user-friendly API built on top of Unreal Engine, allowing researchers to deploy and control LLM-driven agents using natural language instructions. We integrate large-scale LLMs and vision-language models (VLMs), such as GPT-based models, to generate novel environments and structured tasks from multimodal inputs. Our experiments benchmark the performance of several popular LLMs across tasks of increasing complexity, analyzing differences in adaptability, planning, and multi-agent coordination. We also describe our methodology for procedural task generation, task validation, and real-time environment control. VirtualEnv is released as an open-source platform, we aim to advance research at the intersection of AI and gaming, enable standardized evaluation of LLMs in embodied AI settings, and pave the way for future developments in immersive simulations and interactive entertainment.

具身智能仿真平台大模型评估多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。