构建3000+交互任务数据集,评测视觉语言模型在真实环境中的主动推理能力。
EmbRACE-3K: Embodied Reasoning and Action in Complex Environments
- 用虚幻引擎构建3000+多步交互任务,模拟第一人称动态环境
- 零样本下所有模型成功率低于20%,暴露当前模型短板
- 适合作为具身智能研究基准,助力训练更强大交互模型
近期先进视觉语言模型(VLMs)在被动、离线的图像和视频理解任务中表现优异,但在需要在线交互与主动场景理解的具身环境中仍表现有限。在这些场景中,智能体从第一人称视角感知环境,每一步动作都会动态影响后续观测。即使是GPT-4o、Claude 3.5 Sonnet和Gemini 2.5 Pro等顶尖模型,在开放环境交互中也表现出明显局限,尤其在空间推理和长时程规划方面。为填补这一空白,我们提出了EmRACE-3K,一个包含超过3000个语言引导任务的数据集,其环境基于虚幻引擎与UnrealCV-Zoo框架构建,具有高度逼真性。任务涵盖导航、物体操作及多阶段目标执行等多种具身挑战。每个任务以多步轨迹形式展开,配以第一人称视觉观测、高层指令、基础动作及自然语言推理,体现智能体每一步意图。利用EmRACE-3K,我们建立了一个评估VLM具身推理能力的基准,涵盖探索、动态时空语义推理、多阶段目标执行三个维度。在零样本设置下,所有模型成功率均低于20%,凸显本基准的挑战性及当前VLM在交互环境中的局限。为验证数据集效用,我们对Qwen2.5-VL-7B进行监督学习与强化学习微调,显著提升三类任务表现,证明该数据集对推动具身推理能力发展的有效性。
原文摘要 · Abstract (English)
Recent advanced vision-language models(VLMs) have demonstrated strong performance on passive, offline image and video understanding tasks. However, their effectiveness in embodied settings, which require online interaction and active scene understanding remains limited. In such scenarios, an agent perceives the environment from a first-person perspective, with each action dynamically shaping subsequent observations. Even state-of-the-art models such as GPT-4o, Claude 3.5 Sonnet, and Gemini 2.5 Pro struggle in open-environment interactions, exhibiting clear limitations in spatial reasoning and long-horizon planning. To address this gap, we introduce EmRACE-3K, a dataset of over 3,000 language-guided tasks situated in diverse, photorealistic environments constructed using Unreal Engine and the UnrealCV-Zoo framework. The tasks encompass a wide range of embodied challenges, including navigation, object manipulation, and multi-stage goal execution. Each task unfolds as a multi-step trajectory, pairing first-person visual observations with high-level instructions, grounded actions, and natural language rationales that express the agent's intent at every step. Using EmRACE-3K, we establish a benchmark to evaluate the embodied reasoning capabilities of VLMs across three key dimensions: Exploration, Dynamic Spatial-Semantic Reasoning, and Multi-stage Goal Execution. In zero-shot settings, all models achieve success rates below 20%, underscoring the challenge posed by our benchmark and the current limitations of VLMs in interactive environments. To demonstrate the utility of EmRACE-3K, we further fine-tune Qwen2.5-VL-7B using supervised learning followed by reinforcement learning. This approach yields substantial improvements across all three challenge categories, highlighting the dataset's effectiveness in enabling the development of embodied reasoning capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。