提出用残差评估长任务失败原因,揭示真实与预期表现差距。
Benchmarking the Residual: What Long-Horizon Evaluations Add Beyond Matched Short-Task Performance

- 通过对比长任务实际成功率与短阶段预测值,计算'残差'来量化失败程度。
- 发现长任务失败并非仅因错误累积,更多来自上下文积累导致的性能退化。
- 适合评估智能体在复杂任务中稳定性,尤其关注上下文压力下的表现。
长时序基准测试常显示,随着任务变长,智能体失败率上升。这一现象对部署有参考价值,但无法解释失败根源。更多步骤带来更多错误累积机会;长任务可能包含更难的单个决策,或随对话历史、工具输出和环境变化而变得愈发困难。我们用轨迹诱导退化描述后一种情况:早期执行使后期任务更难。当有害累积主要表现为模型可见文本时,常称为上下文腐烂。本文主张,要声称存在‘长时序失败’,必须将完整任务成功率与基于短阶段的基线预测进行比较。我们将该预测值与实际成功率的对数比值称为‘视野残差’。比较需使用相同智能体配置,并预先规定阶段划分、检查点选择、信息传递及资源预算。残差揭示了完整执行流程与基线预测之间的差异,但具体原因仍需针对性实验验证。
原文摘要 · Abstract (English)
Long-horizon benchmarks often show that agents fail more as tasks become longer. This observation is useful for deployment, but it does not by itself explain why failure occurs. More stages create more opportunities for ordinary errors to compound; longer tasks may also contain harder individual decisions or become harder as conversation history, tool outputs, and environment changes accumulate. We use trajectory-induced degradation to mean this last possibility: earlier execution makes later work harder. When the harmful accumulation is specifically the text visible to the model, it is often called context rot. In this position paper, we argue that to claim a "long-horizon failure", benchmarks must compare actual full-task success against a baseline prediction built from short, individual stages. We call the log-ratio between this prediction and actual success the horizon residual. The comparison must use the same agent configuration and specify in advance how stages, checkpoints, information, and budgets will be chosen. The residual shows that the full rollout differs from the chosen baseline; targeted experiments are still needed to explain why.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。