用自然语言指令让大模型自动为小语种做逐词标注,提升标注效率与准确性。
Prompt and circumstance: A word-by-word LLM prompting approach to interlinear glossing for low-resource languages
- 用检索增强的自然语言指令引导大模型逐词标注
- 7种语言中均超越基线,5种语言词级得分超冠军模型
- 能自动生成并遵循语言规则,减少语法歧义错误
部分自动化构建逐词标注文本(IGT)有望助力语言记录。我们认为,大模型因其能理解自然语言指令,可使该过程对语言学家更易用。我们研究了基于检索的大模型提示方法在七种低资源语言上的标注效果,这些语言来自SIGMORPHON 2023共享任务。我们的系统在所有语言的形态词级评分上均优于基于BERT的基线模型;且一个简单的3候选最优方案,在五种语言上的词级得分高于挑战赛优胜者(经过调优的序列模型)。在车兹语案例研究中,我们让大模型自动创建并遵循语言规则,有效降低了复杂语法特征导致的错误。结果表明,大模型在交互式标注系统中具有潜力,既能向人工标注者提供建议,也能准确执行指令。
原文摘要 · Abstract (English)
Partly automated creation of interlinear glossed text (IGT) has the potential to assist in linguistic documentation. We argue that LLMs can make this process more accessible to linguists because of their capacity to follow natural-language instructions. We investigate the effectiveness of a retrieval-based LLM prompting approach to glossing, applied to the seven languages from the SIGMORPHON 2023 shared task. Our system beats the BERT-based shared task baseline for every language in the morpheme-level score category, and we show that a simple 3-best oracle has higher word-level scores than the challenge winner (a tuned sequence model) in five languages. In a case study on Tsez, we ask the LLM to automatically create and follow linguistic instructions, reducing errors on a confusing grammatical feature. Our results thus demonstrate the potential contributions which LLMs can make in interactive systems for glossing, both in making suggestions to human annotators and following directions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。