统一长文本评估标准,高效测试大模型长上下文能力。
LOOM-Scope: a comprehensive and efficient LOng-cOntext Model evaluation framework
- 统一多基准测试的评估设置,确保结果可比性。
- 支持高效推理加速方法部署,降低计算开销。
- 提供轻量全面的评测套件,适合研究与应用验证。
长上下文处理已成为大语言模型的核心能力。尽管已有多个长上下文评估基准提出,但不同基准间的评估设置差异导致结果不一致,难以可靠比较。此外,长上下文评估的高计算成本也制约了社区对长上下文模型的全面评估。本文提出 LOOM-Scope,一个全面且高效的长上下文评估框架。该框架标准化了跨多种基准的评估设置,支持高效长上下文推理加速方法的部署,并引入一套全面且轻量的基准测试套件,实现对模型的综合评估。
原文摘要 · Abstract (English)
Long-context processing has become a fundamental capability for large language models~(LLMs). To assess model's long-context performance, numerous long-context evaluation benchmarks have been proposed. However, variations in evaluation settings across these benchmarks lead to inconsistent results, making it difficult to draw reliable comparisons. Besides, the high computational cost of long-context evaluation poses a significant barrier for the community to conduct comprehensive assessments of long-context models. In this paper, we propose LOOM-Scope, a comprehensive and efficient framework for long-context evaluation. LOOM-Scope standardizes evaluation settings across diverse benchmarks, supports deployment of efficient long-context inference acceleration methods, and introduces a holistic yet lightweight benchmark suite to evaluate models comprehensively. Homepage: https://loomscope.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。