评测大模型在Minecraft中长期探索能力,发现复杂任务下模型表现大幅下降。
MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft

- 构建多智能体合成流程生成可靠开放世界任务实例
- 大模型在长轨迹多跳任务中性能显著下降,需协调隐藏前置条件
- 适合研究具身智能、长程推理与多模态模型评估的学者
多模态大模型在感知、推理和行动生成方面表现强劲,但在动态开放世界中的持续探索能力仍不明确。现有具身与游戏基准常将交互压缩为短时任务,或混杂成功判断与特定游戏机制。本文提出MineExplorer基准,用于评估大模型代理在Minecraft中的开放世界探索能力。首先筛选出依赖Minecraft特有知识的原子任务,以更真实反映通用开放世界推理。随后基于ReAct式能力框架,将原子任务组合为隐含多跳任务。为构建可靠实例,采用多智能体合成工作流,联合设计任务图、沙盒场景与基于规则的里程碑评估器。人类评估显示,该流程生成的实例可靠性显著优于单智能体基线。实验表明,尽管先进大模型能处理多数单跳任务,但在需跨长轨迹协调隐藏前置条件的多跳任务中性能急剧下降。进一步分析发现任务难度与代理完成度正相关,而更大模型或思考模式并不稳定提升表现。代码与数据集已开源。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have shown strong capabilities in perception, reasoning, and action generation. However, their ability to sustain exploration in dynamic open worlds remains unclear. Existing embodied and game-based benchmarks often compress interaction into short-horizon tasks or entangle success with domain-specific game mechanics. In this paper, we introduce MineExplorer benchmark for evaluating open-world exploration capabilities of MLLM agents in Minecraft. We first filter atomic tasks whose solutions rely heavily on Minecraft-specific knowledge to better reflect general open-world reasoning. Then we organize the benchmark around a ReAct-style capability formulation and compose atomic tasks into implicit multi-hop tasks. To further construct reliable instances, MineExplorer uses a multi-agent synthesis workflow that jointly designs task graphs, sandbox scenes, and rule-based milestone evaluators. Human evaluation shows that the multi-agent synthesis workflow produces significantly more reliable instances than a single-agent baseline. Experiments with advanced MLLM agents show that open-world exploration remains challenging, as strong models can handle many single-hop tasks but degrade sharply when hidden prerequisites must be coordinated over longer trajectories. Further analysis finds that task difficulty tracks agent completion, and larger models or thinking modes do not consistently translate into better performance. Code and dataset are available at https://github.com/meituanlongcat/MineExplorer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。