提出新评估方法,检验大模型代理对环境的真实理解能力。
What Do LLM Agents Know About Their World? Task2Quiz: A Paradigm for Studying Environment Understanding
- 设计任务转问答的自动化评测框架,分离任务执行与环境认知。
- 实验发现任务成功不代表环境理解,现有记忆机制效果有限。
- 适合研究通用智能体、自主系统泛化能力的学者参考。
大型语言模型(LLM)代理在复杂决策与工具使用任务中表现出色,但其在不同环境间的泛化能力仍缺乏深入研究。当前评估多依赖轨迹指标衡量任务完成度,却无法检验代理是否具备对环境的稳固、可迁移的认知模型。为此,我们提出任务转问答(T2Q)评估范式,实现确定性与自动化,旨在解耦任务执行与世界状态理解。我们在 T2QBench 中构建了包含30个环境和1,967个基于环境的问答对的评测套件,覆盖多个难度层级。大量实验表明,任务成功率常不能反映真实的环境理解能力,且现有记忆机制难以帮助代理建立稳固的环境认知模型。研究揭示主动探索与细粒度状态表征是主要瓶颈,为构建更具泛化能力的自主智能体提供了坚实基础。
原文摘要 · Abstract (English)
Large language model (LLM) agents have demonstrated remarkable capabilities in complex decision-making and tool-use tasks, yet their ability to generalize across varying environments remains a under-examined concern. Current evaluation paradigms predominantly rely on trajectory-based metrics that measure task success, while failing to assess whether agents possess a grounded, transferable model of the environment. To address this gap, we propose Task-to-Quiz (T2Q), a deterministic and automated evaluation paradigm designed to decouple task execution from world-state understanding. We instantiate this paradigm in T2QBench, a suite comprising 30 environments and 1,967 grounded QA pairs across multiple difficulty levels. Our extensive experiments reveal that task success is often a poor proxy for environment understanding, and that current memory machanism can not effectively help agents acquire a grounded model of the environment. These findings identify proactive exploration and fine-grained state representation as primary bottlenecks, offering a robust foundation for developing more generalizable autonomous agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。