用大模型与规则结合,实现无平行数据的希腊方言转标准语
Dialect Normalization using Large Language Models and Morphological Rules
- 融合语言学规则与大模型的少样本提示方法
- 在希腊方言谚语数据集上取得可比人工标注效果
- 适合低资源方言处理及下游语义分析研究者
自然语言理解系统在低资源语言(包括高资源语言的多种方言)上表现不佳。方言转标准语任务旨在将方言文本转换为标准语言格式,以便下游工具使用。本文提出一种新方法,结合基于语言学的规则变换与大语言模型(LLMs)的针对性少样本提示,无需任何平行语料。我们在希腊方言上实现该方法,并在地区谚语数据集上评估输出质量,由人工标注员打分。随后利用该数据集开展下游实验,发现以往对这些谚语的研究仅依赖表层语言信息(如拼写特征),而经过规范化后仍可挖掘深层语义。
原文摘要 · Abstract (English)
Natural language understanding systems struggle with low-resource languages, including many dialects of high-resource ones. Dialect-to-standard normalization attempts to tackle this issue by transforming dialectal text so that it can be used by standard-language tools downstream. In this study, we tackle this task by introducing a new normalization method that combines rule-based linguistically informed transformations and large language models (LLMs) with targeted few-shot prompting, without requiring any parallel data. We implement our method for Greek dialects and apply it on a dataset of regional proverbs, evaluating the outputs using human annotators. We then use this dataset to conduct downstream experiments, finding that previous results regarding these proverbs relied solely on superficial linguistic information, including orthographic artifacts, while new observations can still be made through the remaining semantics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。