arXiv:2605.28079cs.CL2026-05被引 2

ATLAS为长文本模型提供多尺度能力评估,揭示单一分数的误导性。

ATLAS: All-round Testing of Long-context Abilities across Scales

论文配图:ATLAS: All-round Testing of Long-context Abilities across Scales
图 1 · 摘自论文原文
  • 构建分层评估框架,区分基础操作与应用任务,定位性能失效原因。
  • 采用长度感知AUC评分,完整呈现从8K到100万词元的性能衰减曲线。
  • 提出ATLAScore综合指标,突出模型在不同长度下的能力平衡性,适合研究者和工程师参考。

当前长上下文语言模型声称支持高达百万词元的上下文窗口,但评估通常仅报告单一长度或狭窄任务类别,掩盖了两种失败模式:性能随长度增长而崩溃,以及强检索能力未必能转化为下游应用。我们提出ATLAS,一个重新定义长上下文评估的基准框架,将其转化为长度依赖的能力画像。ATLAS包含三项方法论原则:(i) 分层分类法,将基础操作与应用工作负载分离,便于归因失败;(ii) 长度感知AUC评分,基于8K–1M固定网格对得分-长度曲线积分,取代单点指标,呈现完整退化轨迹;(iii) ATLAScore,对分类维度取调和平均,惩罚能力分布不均,并通过非线性聚合实现端到端不确定性传播。我们在八个能力维度上构建九个可审计组件,共6,438个实例,评估26个模型。Gemini-3.1-Pro-Preview在128K时领先,Claude-Opus-4.6在1M时领先。在ATLASscore@8K-128K与ATLASscore@8K-1M之间,7个模型排名变动至少两位,两个分类层级间仅共享61%的跨模型方差,个别模型排名差距达12位。结果表明,应按能力和长度维度报告长上下文质量,而非依赖单一标题分数。

原文摘要 · Abstract (English)

Long-context language models now advertise context windows up to millions of tokens, yet evaluations typically report a single length or a narrow task family, masking two failure modes: performance can collapse as length grows, and strong retrieval need not transfer to downstream use. We present ATLAS, a benchmarking framework that redefines long-context evaluation as length-dependent capability profiling. ATLAS contributes three methodological principles:(i) a layered taxonomy separating foundational operations from application workloads so failures can be attributed, (ii) length-aware AUC scoring that integrates score-length curves over a fixed 8K-1M grid, replacing single-point metrics with full degradation profiles, and (iii) ATLAScore, a harmonic-mean aggregate over taxonomy categories that penalizes imbalanced profiles, with end-to-end uncertainty propagation from subset scores through the nonlinear final aggregate. We instantiate the framework across eight capability dimensions with nine auditable components and 6,438 instances, and evaluate 26 models. Gemini-3.1-Pro-Preview leads at 128K, Claude-Opus-4.6 leads at 1M. Rankings reshuffle substantially between ATLASscore@8K-128K and ATLASscore@8K-1M: 7 models move by at least two ranks, and the two taxonomy layers share only 61% of cross-model variance, with individual rank gaps up to 12 positions. These results support reporting long-context quality by capability and length, not by a single headline score.

长上下文评估框架模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。