构建了一个覆盖13项阅读理解能力的综合评测基准
MRCEval: A Comprehensive, Challenging and Accessible Machine Reading Comprehension Benchmark
- 基于新分类体系,用大模型生成和筛选题目
- 包含2100道高质量多选题,覆盖13种理解技能
- 适合评估大模型阅读理解能力,尤其关注短板
机器阅读理解(MRC)是评估自然语言理解能力的重要任务。现有MRC数据集主要侧重特定方面,缺乏全面性。为此,我们提出一种新的分类体系,系统划分阅读理解所需的核心能力。基于该体系,构建了MRCEval——一个利用大语言模型(LLMs)作为样本生成器和评判者的新基准。该基准具有综合性、挑战性和可访问性,涵盖13类阅读理解技能,包含2.1K道高质量多选题。我们对28个主流开源与专有模型进行了全面评估,结果显示即使在大模型时代,阅读理解仍面临显著挑战。
原文摘要 · Abstract (English)
Machine Reading Comprehension (MRC) is an essential task in evaluating natural language understanding. Existing MRC datasets primarily assess specific aspects of reading comprehension (RC), lacking a comprehensive MRC benchmark. To fill this gap, we first introduce a novel taxonomy that categorizes the key capabilities required for RC. Based on this taxonomy, we construct MRCEval, an MRC benchmark that leverages advanced Large Language Models (LLMs) as both sample generators and selection judges. MRCEval is a comprehensive, challenging and accessible benchmark designed to assess the RC capabilities of LLMs thoroughly, covering 13 distinct RC skills with a total of 2.1K high-quality multi-choice questions. We perform an extensive evaluation of 28 widely used open-source and proprietary models, highlighting that MRC continues to present significant challenges even in the era of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。