arXiv:2602.05843cs.CL2026-02被引 8

评测大模型在长周期主动探索中的发现能力,揭示其自主推理短板。

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions

  • 构建长周期主动交互的评测框架,强调从经验中归纳规律
  • 120个任务验证模型在长期探索中的发现效率,顶尖模型仍表现不足
  • 适合研究自主智能体、具身智能与复杂环境决策的研究者

大语言模型的发展推动了能够自主导航复杂环境的智能体诞生。然而,现有评估多采用演绎范式,即基于明确规则和静态目标执行任务,通常局限于短周期规划。这忽视了智能体需从经验中自主发现潜在转换规律的归纳必要性,而这是实现智能体远见和维持战略一致性的基础。为此,我们提出OdysseyArena,将智能体评估重新聚焦于长周期、主动且归纳性的交互。我们形式化并实例化四种基本机制,将抽象的转换动态转化为具体交互环境。在此基础上,建立OdysseyArena-Lite用于标准化评测,提供120个任务以衡量智能体的归纳效率与长周期发现能力。进一步地,引入OdysseyArena-Challenge,通过极端交互周期(如超过200步)测试智能体稳定性。对15+领先大模型的实验表明,即使前沿模型在归纳场景中仍显不足,暴露出复杂环境中自主发现的关键瓶颈。代码与数据公开于https://github.com/xufangzhi/Odyssey-Arena。

原文摘要 · Abstract (English)

The rapid advancement of Large Language Models (LLMs) has catalyzed the development of autonomous agents capable of navigating complex environments. However, existing evaluations primarily adopt a deductive paradigm, where agents execute tasks based on explicitly provided rules and static goals, often within limited planning horizons. Crucially, this neglects the inductive necessity for agents to discover latent transition laws from experience autonomously, which is the cornerstone for enabling agentic foresight and sustaining strategic coherence. To bridge this gap, we introduce OdysseyArena, which re-centers agent evaluation on long-horizon, active, and inductive interactions. We formalize and instantiate four primitives, translating abstract transition dynamics into concrete interactive environments. Building upon this, we establish OdysseyArena-Lite for standardized benchmarking, providing a set of 120 tasks to measure an agent's inductive efficiency and long-horizon discovery. Pushing further, we introduce OdysseyArena-Challenge to stress-test agent stability across extreme interaction horizons (e.g., > 200 steps). Extensive experiments on 15+ leading LLMs reveal that even frontier models exhibit a deficiency in inductive scenarios, identifying a critical bottleneck in the pursuit of autonomous discovery in complex environments. Our code and data are available at https://github.com/xufangzhi/Odyssey-Arena

智能体评测长序列决策归纳推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。