用统一语义表示提升GPT-4翻译原住民语言的能力
Can Uniform Meaning Representation Help GPT-4 Translate from Indigenous Languages?
- 在提示中加入统一语义表示(UMR)来辅助翻译
- 多数情况下翻译准确率显著提高,尤其在无示例时
- 适合研究低资源语言与跨语言语义建模的学者
尽管ChatGPT和基于GPT的模型无需微调即可完成多项任务,但在处理极低资源语言和原住民语言时表现不佳。统一语义表示(UMR)是一种旨在捕捉多语言文本语义的语义表征,适用于低资源语言技术开发。本文探索了将UMR融入GPT-4提示以提升其在低资源语言中的下游性能。具体考察了GPT-4在三种原住民语言(纳瓦霍语、阿帕霍语、库卡玛语)上的翻译能力,对比有无示例及有无UMR标注的情况。结果表明,在大多数测试案例中,加入UMR提示可带来统计上显著的性能提升,显示出该形式在未来的广泛应用前景。
原文摘要 · Abstract (English)
While ChatGPT and GPT-based models are able to effectively perform many tasks without additional fine-tuning, they struggle with tasks related to extremely low-resource languages and indigenous languages. Uniform Meaning Representation (UMR), a semantic representation designed to capture the meaning of texts in many languages, is well-positioned to be leveraged in the development of low-resource language technologies. In this work, we explore the downstream utility of UMR for low-resource languages by incorporating it into GPT-4 prompts. Specifically, we examine the ability of GPT-4 to perform translation from three indigenous languages (Navajo, Arápaho, and Kukama), with and without demonstrations, as well as with and without UMR annotations. Ultimately, we find that in the majority of our test cases, integrating UMR into the prompt results in a statistically significant increase in performance, which is a promising indication of future applications of the UMR formalism.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。