构建可调控难度的诊断环境,测试大模型用工具推理的精准性。
ZEBRAARENA: A Diagnostic Simulation Environment for Studying Reasoning-Action Coupling in Tool-Augmented LLMs
- 通过程序化生成环境,分离推理与外部动作的耦合关系。
- 前沿模型如GPT-5在难题上仅达60%准确率,且工具调用效率低。
- 适合研究工具增强型大模型的推理-行动协同机制。
工具增强型大语言模型需将多步推理与外部操作紧密结合,但现有基准常混杂复杂环境动态、记忆知识或数据污染。本文提出ZebraArena,一个程序化生成的诊断环境,用于研究工具增强型大模型的推理-行动耦合,具备可控难度和知识最小化设计,有效避免记忆或数据污染带来的收益。每个任务均需通过特定工具获取关键信息,实现外部信息获取与演绎推理之间的可解释接口。该设计支持确定性评估及理论最优查询次数,用于衡量工具使用的效率。实验表明,ZebraArena要求深度推理与精准工具调用并重,而前沿模型如GPT-5和Gemini 2.5 Pro在难题上仅达60%准确率。此外,实际工具调用次数比理论最优高出70%-270%。我们总结关键发现,希望推动对内部推理与外部动作交互机制的研究。
原文摘要 · Abstract (English)
Tool-augmented large language models (LLMs) must tightly couple multi-step reasoning with external actions, yet existing benchmarks often confound this interplay with complex environment dynamics, memorized knowledge or dataset contamination. In this paper, we introduce ZebraArena, a procedurally generated diagnostic environment for studying reasoning-action coupling in tool-augmented LLMs, with controllable difficulty and a knowledge-minimal design, which limits gains from memorization or dataset contamination. Each task in ZebraArena requires a set of critical information which is available only through targeted tool use, yielding an interpretable interface between external information acquisition and deductive reasoning. This design provides deterministic evaluation via unique solutions, and a theoretical optimal query count for measuring efficient tool use. We show that ZebraArena requires a combination of in-depth reasoning and accurate external tool calling, which remains a challenge as frontier reasoning models such as GPT-5 and Gemini 2.5 Pro only achieves 60% accuracy on the hard instances. We also observe a persistent gaps between theoretical optimality and practical tool usage. For example, GPT-5 uses 70-270% more tool calls than the theoretical optimum. We highlight the key findings in our evaluation, and hope ZebraArena stimulates further research on the interplay between internal reasoning and external action.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。