arXiv:2604.17108cs.CLcs.AI2026-04ACL

首个希伯来语指代消解基准,解决复杂词形带来的边界难题。

Beyond Word Boundaries: A Hebrew Coreference Benchmark and an Evaluation Protocol for Morphologically Complex Text

  • 构建多层级标注的希伯来语指代数据集,支持词、子词与多词提及
  • 大模型在未分词文本上表现显著下降,小模型反而更优
  • 提出考虑词素边界的评估协议,适配形态丰富的语言研究

指代消解(CR)是信息抽取、摘要等长文本任务的关键,但现有方法多基于英语设计,难以应对形态丰富的语言(MRLs)。在这些语言中,提及边界常不等于词边界,单个词可能包含多个回指成分。当前主流模型和评估协议默认词与提及对齐,这一假设在希伯来语等语言中失效。为此,我们推出首个希伯来语核心指代数据集KibutzR,涵盖词、子词及多词层级的提及标注,并提出针对词/词素边界差异的评估协议。实验表明,当代大模型在希伯来语上的表现显著低于英语,且在原始未分词文本上性能进一步下降;值得注意的是,小型编码器模型反而优于主流解码器模型,展现出反向性能趋势。本工作为希伯来语指代消解提供新基准与可推广的评估框架,推动其他形态复杂语言的研究。

原文摘要 · Abstract (English)

Coreference Resolution (CR) is a fundamental NLP task critical for long-form tasks as information extraction, summarization, and many business applications. However, CR methods originally designed for English struggle with Morphologically Rich Languages (MRLs), where mention boundaries do not necessarily align with word boundaries, and a single token may consist of multiple anaphors. CR modeling and evaluation protocols standardly assume that, as in English, words and mentions mostly align. However, this assumption breaks down in MRLs, particularly in the context of LLMs' raw-text processing and end-to-end tasks. To assess and address this challenge, we introduce {\em KibutzR}, the first comprehensive CR dataset for Modern Hebrew, an MRL rich with complex words and pronominal clitics. We deliver an annotated dataset that identifies mentions at word, sub-word and multi-word levels, and propose an evaluation protocol that directly addresses word/morpheme boundary discrepancies. Our experiments show that contemporary LLMs perform significantly worse on Hebrew than on English, and that performance degrades on raw unsegmented text. Crucially, we show an inverse performance-trend in Hebrew relative to English, where smaller encoders perform far better than contemporary decoder models, leaving ample space for investigation and improvement. We deliver a new benchmark for Hebrew coreference resolution and a segmentation-aware evaluation protocol to inform future work on other MRLs.

指代消解形态丰富语言多语言NLP基准数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。