用AI抢救濒危汉字——仅35例样本实现48.69%翻译准确率
NushuRescue: Revitalization of the Endangered Nushu Language with AI
- 基于少量样本和GPT-4-Turbo构建自动训练框架
- 仅用35个示例达成48.69%翻译准确率,生成98条新语料
- 适合语言保护、低资源语言研究者使用
濒危语言的保护与复兴对文化传承及语言学、人类学研究具有重要意义,但这些语言普遍资源匮乏,重建成本高昂。以中国瑶族女性曾使用的独特文字Nushu为例,本文提出NushuRescue——一种面向低资源濒危语言的大语言模型训练框架。该框架通过自动化评估与语料扩展,加速语言复兴进程。作为基础数据集,我们构建了首个公开的500句Nushu-中文平行语料库NCGold。利用仅35个短例的NCGold,GPT-4-Turbo在50条保留句子上实现48.69%的翻译准确率,并生成98条新的现代汉语译文(称作NCSilver)。此外,还开发了FastText与Seq2Seq模型支持研究。NushuRescue为濒危语言复兴提供可扩展、低人力依赖的解决方案。
原文摘要 · Abstract (English)
The preservation and revitalization of endangered and extinct languages is a meaningful endeavor, conserving cultural heritage while enriching fields like linguistics and anthropology. However, these languages are typically low-resource, making their reconstruction labor-intensive and costly. This challenge is exemplified by Nushu, a rare script historically used by Yao women in China for self-expression within a patriarchal society. To address this challenge, we introduce NushuRescue, an AI-driven framework designed to train large language models (LLMs) on endangered languages with minimal data. NushuRescue automates evaluation and expands target corpora to accelerate linguistic revitalization. As a foundational component, we developed NCGold, a 500-sentence Nushu-Chinese parallel corpus, the first publicly available dataset of its kind. Leveraging GPT-4-Turbo, with no prior exposure to Nushu and only 35 short examples from NCGold, NushuRescue achieved 48.69% translation accuracy on 50 withheld sentences and generated NCSilver, a set of 98 newly translated modern Chinese sentences of varying lengths. A sample of both NCGold and NCSilver is included in the Supplementary Materials. Additionally, we developed FastText-based and Seq2Seq models to further support research on Nushu. NushuRescue provides a versatile and scalable tool for the revitalization of endangered languages, minimizing the need for extensive human input.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。