arXiv:2606.13218cs.CL2026-06中稿 · EMNLP

测试大模型对阿拉伯语希伯来语同源词的语义区分能力

When Similar Means Different: Evaluating LLMs on Arabic--Hebrew Cognates

论文配图:When Similar Means Different: Evaluating LLMs on Arabic--Hebrew Cognates
图 1 · 摘自论文原文
  • 构建1858组阿拉伯-希伯来词对,带句子级标注的语义辨析数据集
  • 模型普遍依赖表面形式相似性,对伪同源词识别差,借词表现波动大
  • 原文字形输入效果最好,上下文和规模提升因模型而异

阿拉伯语与希伯来语作为密切相关的闪米特语系语言,共享大量表面形式相似的词汇,包括真正的同源词、伪同源词以及现代借词。这些词汇关系为跨语言语义理解带来歧义:形式相似的词可能对应相同、不同或借用的含义。为评估大模型能否做出此类区分,我们引入SemCog Bench,一个包含1858组阿拉伯-希伯来词对的精选基准,附有句子级别的同源词识别与语义消歧标注。我们在多种输入表示和上下文设置下评估了多类开源与专有大模型。结果显示模型高度依赖表面形式相似性,对伪同源词表现较弱,对借词的表现差异显著。进一步发现,上下文和模型规模带来的提升具有模型依赖性,原文字形输入通常表现最佳。研究揭示了大模型在跨语言形式-语义推理上的局限性,并确立SemCog Bench作为跨语言词汇语义研究的基准。代码与数据已公开。

原文摘要 · Abstract (English)

Arabic and Hebrew, as closely related Semitic languages, share many words with similar surface forms, including true cognates, false friends, and modern loanwords. These lexical relationships create ambiguity for cross-lingual semantic interpretation, as similar forms may correspond to shared, divergent, or borrowed meanings. To evaluate whether LLMs can make these distinctions, we introduce SemCog Bench, a curated benchmark of 1,858 Arabic--Hebrew word pairs with sentence-level annotations for cognate identification and semantic disambiguation. We evaluate a diverse set of open-source and proprietary LLMs across multiple input representations and contextual settings. Our results show reliance on surface-form similarity, with weaker performance on false friends and wide variation on loanwords. We further find that context and scale yield model-dependent gains, while original-script inputs generally perform best. Our findings reveal limitations in cross-lingual form-meaning reasoning and establish SemCog Bench as a benchmark for cross-lingual lexical semantics. Our code and data are publicly available.

大模型评估跨语言语义消歧阿拉伯语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。