arXiv:2601.20730cs.CL2026-01被引 7

用环境演算测试智能体长上下文能力,发现动态信息整合才是难点。

AgentLongBench: A Controllable Long Benchmark For Long-Contexts Agents via Environment Rollouts

  • 基于横向思维谜题生成动态环境轨迹,模拟真实交互
  • 32K到400万token下,模型在动态推理中性能显著下降
  • 信息密度高比记忆碎片化更难处理,适合研究长程智能体的学者

大型语言模型向自主智能体演进,亟需管理长而动态的上下文。现有基准多为静态,依赖被动检索任务,无法模拟智能体与环境互动中的非线性推理和迭代反馈等复杂性。为此,我们提出 extbf{AgentLongBench},通过基于横向思维谜题的环境演算来评估智能体。该框架在知识密集型与知识无关场景下生成严格的交互轨迹。对前沿模型与记忆系统(支持32K至400万token)的实验表明:尽管在静态检索中表现良好,智能体在工作流所需的动态信息整合上存在明显短板。分析显示,这一退化主要由解决查询所需的最小令牌数决定。该因素解释了为何大规模工具响应的信息密度远比长对话中的记忆碎片化更具挑战性。

原文摘要 · Abstract (English)

The evolution of Large Language Models (LLMs) into autonomous agents necessitates the management of extensive, dynamic contexts. Current benchmarks, however, remain largely static, relying on passive retrieval tasks that fail to simulate the complexities of agent-environment interaction, such as non-linear reasoning and iterative feedback. To address this, we introduce \textbf{AgentLongBench}, which evaluates agents through simulated environment rollouts based on Lateral Thinking Puzzles. This framework generates rigorous interaction trajectories across knowledge-intensive and knowledge-free scenarios. Experiments with state-of-the-art models and memory systems (32K to 4M tokens) expose a critical weakness: while adept at static retrieval, agents struggle with the dynamic information synthesis essential for workflows. Our analysis indicates that this degradation is driven by the minimum number of tokens required to resolve a query. This factor explains why the high information density inherent in massive tool responses poses a significantly greater challenge than the memory fragmentation typical of long-turn dialogues.

长上下文智能体评估动态推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。