arXiv:2501.17569cs.CL2025-01EMNLP被引 3

用语言学特征评估阅读理解模型短板,发现大模型仍会因复杂句式失败。

A linguistically-motivated evaluation methodology for unraveling model's abilities in reading comprehension tasks

  • 基于语义框架标注语言复杂度,提取七类影响因素
  • 法语数据集验证:两类复杂度是模型失败的关键预测因子
  • 用Chat-GPT辅助标注英语数据,揭示当前模型能力盲区

我们提出一种基于语言学直觉的阅读理解评估方法,认为某些例题因其语言复杂性,无论模型大小或结构如何,都会持续导致较低得分。通过语义框架标注刻画这种复杂性,并研究了七个可能影响模型表现的复杂度因素。首先在精心标注的法语阅读理解基准上应用该方法,发现其中两个复杂度因素确实是模型失败的良好预测指标,而其他因素作用较弱。随后,我们在一个广为人知的英文基准上使用Chat-GPT作为语义标注代理,进一步部署该方法。研究结果表明,对阅读理解任务进行细粒度的语言学驱动自动评估不仅可行,而且有助于理解模型处理输入中特定语言特征的能力。同时显示,当前最先进的模型在处理某些语言特征时依然存在明显缺陷,说明仅增加模型规模不足以有效应对这些挑战。

原文摘要 · Abstract (English)

We introduce an evaluation methodology for reading comprehension tasks based on the intuition that certain examples, by the virtue of their linguistic complexity, consistently yield lower scores regardless of model size or architecture. We capitalize on semantic frame annotation for characterizing this complexity, and study seven complexity factors that may account for model's difficulty. We first deploy this methodology on a carefully annotated French reading comprehension benchmark showing that two of those complexity factors are indeed good predictors of models' failure, while others are less so. We further deploy our methodology on a well studied English benchmark by using Chat-GPT as a proxy for semantic annotation. Our study reveals that fine-grained linguisticallymotivated automatic evaluation of a reading comprehension task is not only possible, but helps understand models' abilities to handle specific linguistic characteristics of input examples. It also shows that current state-of-the-art models fail with some for those characteristics which suggests that adequately handling them requires more than merely increasing model size.

阅读理解语言学评估方法模型缺陷

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。