用论文引用结构自动构建学术长文本推理评测基准
SCALAR: Scientific Citation-based Live Assessment of Long-context Academic Reasoning
- 基于论文引用关系自动生成无人工标注的评测标签
- 人类专家准确率超90%,多数模型在闭合题中不足50%
- 支持动态更新,适合追踪学术推理能力进展
长上下文理解已成为大语言模型的关键能力,但其评估仍具挑战。我们提出SCALAR,一个用于评估学术写作中基于引用的长上下文推理能力的基准。SCALAR利用学术论文及其引用结构,自动生成高质量的真值标签,无需人工标注。该基准具备可调节难度和动态更新机制,有效避免数据污染。包含两项任务:多选问答和填空式引用预测。我们评估了多种先进大模型,发现多选任务能有效区分模型能力;人类专家准确率超过90%,而多数模型表现不佳。填空任务更具挑战性,无一模型准确率超过50%。SCALAR提供了一个领域贴合、持续更新的框架,用于跟踪基于引用的长上下文理解进展。
原文摘要 · Abstract (English)
Long-context understanding has emerged as a critical capability for large language models (LLMs). However, evaluating this ability remains challenging. We present SCALAR, a benchmark designed to assess citation-grounded long-context reasoning in academic writing. SCALAR leverages academic papers and their citation structure to automatically generate high-quality ground-truth labels without human annotation. It features controllable difficulty levels and a dynamic updating mechanism that mitigates data contamination. The benchmark includes two tasks: a multiple-choice QA format and a cloze-style citation prediction. We evaluate a range of state-of-the-art LLMs and find that the multiple-choice task effectively distinguishes model capabilities. While human experts achieve over 90% accuracy, most models struggle. The cloze-style task is even more challenging, with no model exceeding 50% accuracy. SCALAR provides a domain-grounded, continuously updating framework for tracking progress in citation-based long-context understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。