arXiv:2508.01812cs.CLcs.AI2025-08EMNLP被引 15

构建首个希伯来语阅读理解基准,解决形态复杂导致的标注难题。

HeQ: a Large and Diverse Hebrew Reading Comprehension Benchmark

  • 设计新标注指南与众包流程,应对希伯来语形态复杂性。
  • 发布包含30147个问答对的HeQ数据集,覆盖维基与新闻文本。
  • 发现传统评估指标不适用于希伯来语,提出改进方案,适合形态丰富语言。

当前希伯来语自然语言处理基准主要聚焦形态句法任务,忽视语义理解维度。为弥补这一缺口,我们构建了一个希伯来语机器阅读理解(MRC)数据集,将MRC定义为抽取式问答。希伯来语形态丰富,复杂形式导致答案跨度边界模糊、标注不一致,进而影响标准评估指标。为此,我们制定新型标注指南、受控众包协议及适配形态丰富语言的评估指标。最终发布的基准HeQ(Hebrew QA)包含30,147个来自希伯来语维基百科和以色列科技新闻的多样化问答对。实证分析表明,标准指标如F1和精确匹配(EM)不适用于希伯来语(及其他形态丰富语言),并提出相应改进。此外,实验显示模型在形态句法任务与阅读理解任务上的表现相关性低,提示专用于前者的模型可能在语义任务上表现不佳。HeQ的开发与探索揭示了形态丰富语言在自然语言理解中的挑战,推动希伯来语及其他类似语言的更好理解模型发展。

原文摘要 · Abstract (English)

Current benchmarks for Hebrew Natural Language Processing (NLP) focus mainly on morpho-syntactic tasks, neglecting the semantic dimension of language understanding. To bridge this gap, we set out to deliver a Hebrew Machine Reading Comprehension (MRC) dataset, where MRC is to be realized as extractive Question Answering. The morphologically rich nature of Hebrew poses a challenge to this endeavor: the indeterminacy and non-transparency of span boundaries in morphologically complex forms lead to annotation inconsistencies, disagreements, and flaws in standard evaluation metrics. To remedy this, we devise a novel set of guidelines, a controlled crowdsourcing protocol, and revised evaluation metrics that are suitable for the morphologically rich nature of the language. Our resulting benchmark, HeQ (Hebrew QA), features 30,147 diverse question-answer pairs derived from both Hebrew Wikipedia articles and Israeli tech news. Our empirical investigation reveals that standard evaluation metrics such as F1 scores and Exact Match (EM) are not appropriate for Hebrew (and other MRLs), and we propose a relevant enhancement. In addition, our experiments show low correlation between models' performance on morpho-syntactic tasks and on MRC, which suggests that models designed for the former might underperform on semantics-heavy tasks. The development and exploration of HeQ illustrate some of the challenges MRLs pose in natural language understanding (NLU), fostering progression towards more and better NLU models for Hebrew and other MRLs.

阅读理解希伯来语数据集形态学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。