arXiv:2603.29231cs.AI2026-03被引 5

提出可靠性评估框架,揭示大模型长时任务表现的隐性风险。

Beyond pass@1: A Reliability Science Framework for Long-Horizon LLM Agents

  • 构建四维可靠性指标体系,量化模型长期任务稳定性
  • 发现高能力模型反而更易崩溃,长时任务排名与能力排名严重背离
  • 内存支架普遍降低长期性能,适合关注系统稳定性的研发者参考

现有基准仅衡量单次成功能力,但实际部署需可靠持续表现。我们发现任务时长增长时,能力与可靠性系统性偏离,且pass@1在短任务上无法察觉此差异。为此提出长时序大模型代理的可靠性科学框架,包含四个指标:可靠性衰减曲线(RDC)、方差放大因子(VAF)、渐进退化得分(GDS)和崩塌起始点(MOP)。在涵盖4个时长段、3个领域共396个任务的基准上,对10个模型进行了23,392次实验。关键发现:(1) 可靠性衰减具有领域特征——搜索增强类任务的GDS从0.90降至0.44,文档处理几乎不变(0.74→0.71);(2) VAF按能力层级分叉——高VAF反映能力强,非不稳定信号;(3) 能力与可靠性排名显著分化,长时任务出现多级倒置;(4) 前沿模型崩塌率最高(达19%),因其尝试复杂多步策略易失控;(5) 所有10个模型中,记忆支架均普遍损害长期表现。结果表明,可靠性应作为与能力并列的一级评估维度。

原文摘要 · Abstract (English)

Existing benchmarks measure capability -- whether a model succeeds on a single attempt -- but production deployments require reliability -- consistent success across repeated attempts on tasks of varying duration. We show these properties diverge systematically as task duration grows, and that pass@1 on short tasks is structurally blind to this divergence. We introduce a reliability science framework for long-horizon LLM agents with four metrics: Reliability Decay Curve (RDC), Variance Amplification Factor (VAF), Graceful Degradation Score (GDS), and Meltdown Onset Point (MOP). We evaluate 10 models across 23,392 episodes on a 396-task benchmark spanning four duration buckets and three domains. Key findings: (1) reliability decay is domain-stratified -- SE GDS drops from 0.90 to 0.44 while document processing is nearly flat (0.74 to 0.71); (2) VAF bifurcates by capability tier -- high VAF is a capability signature, not an instability signal; (3) capability and reliability rankings diverge substantially, with multi-rank inversions at long horizons; (4) frontier models have the highest meltdown rates (up to 19%) because they attempt ambitious multi-step strategies that sometimes spiral; and (5) memory scaffolds universally hurt long-horizon performance across all 10 models. These results motivate reliability as a first-class evaluation dimension alongside capability.

大模型评估可靠性长时任务智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。