构建跨语言学术推理评测基准,检验大模型在复杂科研场景下的理解能力
ScholarBench: A Bilingual Benchmark for Abstraction, Comprehension, and Reasoning Evaluation in Academic Contexts
- 基于学术文献设计五类问题,覆盖八大学科领域
- 包含5031个韩语与5309个英语样本,顶尖模型平均得分仅0.543
- 支持中英双语评估,侧重抽象理解与逻辑推理能力
现有大语言模型领域知识评测基准难以应对复杂的学术任务。为此,我们提出 exttt{ScholarBench},一个聚焦深度专业知识与复杂学术问题求解的基准,通过三步流程构建,评估大模型在学术推理方面的能力。该基准针对来自学术文献的更专业、逻辑更复杂的场景,涵盖五种不同问题类型。与以往基准不同, exttt{ScholarBench} 在八个不同研究领域中评估大模型的抽象、理解与推理能力。为确保数据质量,我们定义了各领域的特定示例属性,并设计与各领域研究方法和论述结构一致的问题。此外,该基准为英韩双语数据集,支持对大模型在两种语言中的语言能力进行同步评估。基准包含5,031个韩语样本和5,309个英语样本,即使最先进的模型如o3-mini,平均得分也仅为0.543,体现出该基准的挑战性。
原文摘要 · Abstract (English)
Prior benchmarks for evaluating the domain-specific knowledge of large language models (LLMs) lack the scalability to handle complex academic tasks. To address this, we introduce \texttt{ScholarBench}, a benchmark centered on deep expert knowledge and complex academic problem-solving, which evaluates the academic reasoning ability of LLMs and is constructed through a three-step process. \texttt{ScholarBench} targets more specialized and logically complex contexts derived from academic literature, encompassing five distinct problem types. Unlike prior benchmarks, \texttt{ScholarBench} evaluates the abstraction, comprehension, and reasoning capabilities of LLMs across eight distinct research domains. To ensure high-quality evaluation data, we define category-specific example attributes and design questions that are aligned with the characteristic research methodologies and discourse structures of each domain. Additionally, this benchmark operates as an English-Korean bilingual dataset, facilitating simultaneous evaluation for linguistic capabilities of LLMs in both languages. The benchmark comprises 5,031 examples in Korean and 5,309 in English, with even state-of-the-art models like o3-mini achieving an average evaluation score of only 0.543, demonstrating the challenging nature of this benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。