破解大模型长任务失败的根源,系统梳理六大关键环节。
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents

- 区分长时序、长上下文、长记忆三类属性,厘清概念混淆。
- 发现任务越长,结果信号越失效,需强化步骤级反馈。
- 适合关注大模型长期任务能力的研究者与工程师。
前沿语言模型可在单次前向传播中解决曾需多年研究的推理问题,但在多小时任务中却频繁失败:遗忘早期决策、过早宣告任务完成或偏离目标。我们称之为‘视野缺口’,并系统调研了1,547篇arXiv论文(2024–2026年),通过种子筛选与26.8%的泄露过滤器收集数据,辅以定向补充。我们明确区分三个常被混淆的属性:长时序(任务所需步数)、长上下文(模型的令牌容量)和长程记忆(系统跨步骤/会话的持续性)。将文献按长时序任务生命周期分为六类:规划、记忆、执行、训练、评估及基础/安全,并交叉分析其信息承载位置(上下文内、任务超上下文、跨任务持久)。在所有类别中,均发现相同模式:随着任务时长增加,仅依赖最终结果的信号变得无用,领域应对策略——如过程奖励模型、信用分配或轨迹诊断——均致力于生成更密集的步骤级信号。我们将批判性与诊断性文献视为核心线索,主张不应将批评与方法分离。最后提出开放测量问题:模型能力与利用能力的解耦、过程信号中相关偏差的管理,以及长时序可靠性是否具备普适预测理论。
原文摘要 · Abstract (English)
Frontier language models solve reasoning problems in a single forward pass that would have been research contributions years ago, yet fail at multi-hour tasks: losing track of earlier decisions, declaring half-finished work done, or drifting from goals. We call this the horizon gap and survey 1,547 arXiv papers (2024-2026) collected via systematic seed harvest with a disclosed 26.8% bleed filter, extended by targeted supplementation. We disambiguate three routinely conflated properties: long-horizon (task property: required steps), long-context (model property: token capacity), and long-term memory (system property: persistence across steps/sessions). We organize the corpus into six categories tracking a long-horizon task's lifecycle -- planning, memory, execution, training, evaluation, and foundations/safety -- crossed with an axis capturing where horizons are carried (within-context, within-task-beyond-context, or cross-task-persistent). Across all categories, we find the same pattern: outcome-only signals grow uninformative as horizons lengthen, and the field's response -- whether process reward models, credit assignment, or trajectory-level diagnostics -- manufactures denser step-level signals. We treat critical and diagnostic literature as first-class threads throughout, arguing that segregating critique from method would routinely split single papers across chapters. We close by naming open measurement problems: decomposing model versus harness capability, managing correlated bias in process-level signals used for both training and evaluation, and whether long-horizon reliability admits general predictive theory.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。