arXiv:2509.17855cs.CL2025-09EMNLP被引 3

用单语语料构建德语-巴伐利亚方言词典,评估大模型对方言词的识别能力。

Make Every Letter Count: Building Dialect Variation Dictionaries from Monolingual Corpora

  • 仅用单语数据构建方言词对标注框架DiaLemma
  • 10万组人工标注词对显示模型在名词上表现最好
  • 上下文提升翻译效果但削弱变体识别能力

方言因缺乏标准拼写而存在显著差异,而大语言模型(LLM)对方言的处理能力尚未被充分研究。本文以巴伐利亚方言为例,通过考察大模型对不同词性方言词的识别与翻译能力,评估其词汇方言理解水平。为此,我们提出一种名为DiaLemma的新标注框架,仅基于单语数据构建方言变异词典,并生成包含10万组人工标注的德语-巴伐利亚词对基准数据集。我们评估了九个前沿大模型在判断巴伐利亚词是否为给定德语词根的方言翻译、屈折变体或无关形式上的表现。结果表明,大模型在名词及词形相似的词对上表现最佳,但在区分直接翻译与屈折变体方面最弱。有趣的是,提供例句上下文可提升翻译性能,却降低模型识别方言变体的能力。该研究揭示了大模型在处理拼写方言差异方面的局限性,强调需进一步改进模型对方言的适应性。

原文摘要 · Abstract (English)

Dialects exhibit a substantial degree of variation due to the lack of a standard orthography. At the same time, the ability of Large Language Models (LLMs) to process dialects remains largely understudied. To address this gap, we use Bavarian as a case study and investigate the lexical dialect understanding capability of LLMs by examining how well they recognize and translate dialectal terms across different parts-of-speech. To this end, we introduce DiaLemma, a novel annotation framework for creating dialect variation dictionaries from monolingual data only, and use it to compile a ground truth dataset consisting of 100K human-annotated German-Bavarian word pairs. We evaluate how well nine state-of-the-art LLMs can judge Bavarian terms as dialect translations, inflected variants, or unrelated forms of a given German lemma. Our results show that LLMs perform best on nouns and lexically similar word pairs, and struggle most in distinguishing between direct translations and inflected variants. Interestingly, providing additional context in the form of example usages improves the translation performance, but reduces their ability to recognize dialect variants. This study highlights the limitations of LLMs in dealing with orthographic dialect variation and emphasizes the need for future work on adapting LLMs to dialects.

方言识别大模型词典构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。