提出可调长度的长文本评估框架,更真实地测试大模型长上下文能力。
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?
- 设计可动态调整输入长度的评测基准,突破固定长度限制。
- 引入新指标分离模型基础能力与真正长上下文处理能力。
- 适用于不同模型对比,能揭示模型失效临界点,适合评测研究者使用。
长上下文能力被视为大语言模型最重要的能力之一,因其能使用户轻松处理原本耗时的任务——例如,从长文档中查找答案,而非直接提问。然而,现有的基于真实任务的长上下文评估基准存在两大缺陷:其一,如LongBench等基准未提供有效指标来区分长上下文表现与模型基线能力,导致跨模型比较模糊;其二,这些基准通常采用固定输入长度,限制了在不同模型间的适用性,并无法揭示模型开始失效的临界点。为此,我们提出一个可调节输入长度的长上下文基准及一种新指标,能够解耦基线知识与真正的长上下文能力。实验表明,该方法在有效评估大语言模型方面具有显著优势。
原文摘要 · Abstract (English)
Long-context capability is considered one of the most important abilities of LLMs, as a truly long context-capable LLM enables users to effortlessly process many originally exhausting tasks -- e.g., digesting a long-form document to find answers vs. directly asking an LLM about it. However, existing real-task-based long-context evaluation benchmarks have two major shortcomings. First, benchmarks like LongBench often do not provide proper metrics to separate long-context performance from the model's baseline ability, making cross-model comparison unclear. Second, such benchmarks are usually constructed with fixed input lengths, which limits their applicability across different models and fails to reveal when a model begins to break down. To address these issues, we introduce a length-controllable long-context benchmark and a novel metric that disentangles baseline knowledge from true long-context capabilities. Experiments demonstrate the superiority of our approach in effectively evaluating LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。