构建首个中英双语长文本评估基准,真实还原复杂场景。
LongBench Pro: A More Realistic and Comprehensive Bilingual Long-Context Evaluation Benchmark
- 用人类与大模型协作生成高质量长文本题目
- 覆盖1.5万样本,最长达256k token,含25项任务
- 适合研究长文本理解与跨语言模型的开发者
大型语言模型上下文长度迅速扩展,但现有评估基准难以跟上。当前长文本评测要么过于合成化,要么人工标注成本过高。我们提出 LongBench Pro,一个包含1,500个自然产生的中英文长文本样本的综合性双语评测集,涵盖11个主任务和25个次级任务,输入长度从8,000到256,000个标记。该基准支持细粒度分析,包含任务特异性指标及多维度上下文需求分类(完整/部分依赖、六级长度、四级难度,由模型表现校准)。为平衡质量与可扩展性,我们提出人-模型协同构建流程:前沿大模型先生成挑战性问题、参考答案及推理过程,降低专家验证成本;专家随后严格验证正确性并修正问题案例。在 LongBench Pro 上评估46个主流长上下文 LLM,发现:(1) 长上下文优化对理解能力的提升优于参数量扩展;(2) 实际有效上下文长度通常短于宣称长度,且存在显著跨语言偏差;(3) “思考”范式主要帮助原生训练过推理能力的模型,而混合思考设计提供了理想的权衡方案。总体而言,LongBench Pro 为推进长上下文理解提供了一个可靠的测试平台。
原文摘要 · Abstract (English)
The rapid expansion of context length in large language models (LLMs) has outpaced existing evaluation benchmarks. Current long-context benchmarks often trade off scalability and realism: synthetic tasks underrepresent real-world complexity, while fully manual annotation is costly to scale to extreme lengths and diverse scenarios. We present LongBench Pro, a more realistic and comprehensive bilingual benchmark of 1,500 naturally occurring long-context samples in English and Chinese spanning 11 primary tasks and 25 secondary tasks, with input lengths from 8k to 256k tokens. LongBench Pro supports fine-grained analysis with task-specific metrics and a multi-dimensional taxonomy of context requirement (full vs. partial dependency), length (six levels), and difficulty (four levels calibrated by model performance). To balance quality with scalability, we propose a Human-Model Collaborative Construction pipeline: frontier LLMs draft challenging questions and reference answers, along with design rationales and solution processes, to reduce the cost of expert verification. Experts then rigorously validate correctness and refine problematic cases. Evaluating 46 widely used long-context LLMs on LongBench Pro yields three findings: (1) long-context optimization contributes more to long-context comprehension than parameter scaling; (2) effective context length is typically shorter than the claimed context length, with pronounced cross-lingual misalignment; and (3) the "thinking" paradigm helps primarily models trained with native reasoning, while mixed-thinking designs offer a promising Pareto trade-off. In summary, LongBench Pro provides a robust testbed for advancing long-context understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。