首个针对美因茨方言的NLP研究,发现大模型无法有效生成或理解该方言。
Meenz bleibt Meenz, but Large Language Models Do Not Speak Its Dialect
- 构建首个可直接用于NLP研究的美因茨方言词典,含2351个词条。
- 大模型生成方言词定义准确率仅6.27%,生成方言词准确率低至1.51%。
- 少样本学习和规则提取仍无法显著提升性能,凸显方言研究资源匮乏。
美因茨方言(Meenzerisch)是德国美因茨市的传统方言,也是当地狂欢节的标志性语言,但正面临消亡危机,与众多德语方言命运相似。自然语言处理(NLP)有望助力方言保护与复兴,但此前尚无专门针对美因茨方言的NLP研究。本文首次开展该领域研究,基于现有资料(Schramm, 1966)构建首个面向NLP的数字词典,包含2,351个美因茨方言词汇及其标准德语释义。利用该数据集,我们探究两大问题:(1) 当前主流大语言模型(LLMs)能否生成方言词的定义?(2) 能否根据标准德语释义生成对应的美因茨方言词?实验表明,最佳模型在定义生成任务中准确率仅为6.27%,词生成任务准确率更低至1.51%。进一步通过少样本学习与从训练集中提取规则输入模型的方式尝试提升效果,虽略有改善,但准确率仍低于10%。结果表明,亟需更多专门资源与研究投入以推动德语方言的数字化保护。
原文摘要 · Abstract (English)
Meenzerisch, the dialect spoken in the German city of Mainz, is also the traditional language of the Mainz carnival, a yearly celebration well known throughout Germany. However, Meenzerisch is on the verge of dying out-a fate it shares with many other German dialects. Natural language processing (NLP) has the potential to help with the preservation and revival efforts of languages and dialects. However, so far no NLP research has looked at Meenzerisch. This work presents the first research in the field of NLP that is explicitly focused on the dialect of Mainz. We introduce a digital dictionary-an NLP-ready dataset derived from an existing resource (Schramm, 1966)-to support researchers in modeling and benchmarking the language. It contains 2,351 words in the dialect paired with their meanings described in Standard German. We then use this dataset to answer the following research questions: (1) Can state-of-the-art large language models (LLMs) generate definitions for dialect words? (2) Can LLMs generate words in Meenzerisch, given their definitions? Our experiments show that LLMs can do neither: the best model for definitions reaches only 6.27% accuracy and the best word generation model's accuracy is 1.51%. We then conduct two additional experiments in order to see if accuracy is improved by few-shot learning and by extracting rules from the training set, which are then passed to the LLM. While those approaches are able to improve the results, accuracy remains below 10%. This highlights that additional resources and an intensification of research efforts focused on German dialects are desperately needed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。