用语法书生成合成数据,提升濒危语言的机器翻译效果。
A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books

- 从语法书中提取规则、例句和词库,自动构建翻译训练数据。
- 在3种语言上测试,最佳情况下翻译准确率提升超8.8分。
- 适合想为资源极少语言开发翻译工具的研究者或保护项目。
多数濒危语言缺乏机器翻译所需的平行语料,尽管已有描述性语法书。我们提出一个流程:利用大语言模型从语法书中提取语法规则、例句和词典,生成可用于微调的合成平行语料——不同于以往在推理时将语法内容注入提示,本方法直接用于训练。在三种类型差异大的低资源语言(卡拉芒语,Papuan;图茨钦语,Romance;曼丹语,Siouan)上验证,结果表明,在96种配置组合中,75%的配置下对卡拉芒语、59%对图茨钦语优于基线种子数据,最佳情况下的ChrF++得分分别提升+8.8、+5.3和+3.3。通过系统性因子实验,我们识别出哪些因素组合带来增益及失效边界。结果证明,静态语言文献可被重用于机器翻译微调,为严重资源匮乏语言提供实用的翻译工具路径。
原文摘要 · Abstract (English)
Most endangered languages lack the parallel data required for machine translation, despite the existence of descriptive grammar books. We introduce a pipeline that uses large language models to extract grammatical rules, example sentences, and lexicons from grammar books and generate synthetic parallel corpora for fine-tuning-rather than feeding grammar content into prompts at inference time, as in prior work. Validated on three typologically diverse low-resource languages-Kalamang (Papuan), Tuatschin (Romance), and Mandan (Siouan)-we show that fine-tuning on synthetic data improves over seed-data baselines in 75% of configurations for Kalamang and 59% for Tuatschin, with best-case ChrF++ gains of +8.8, +5.3, and +3.3 respectively. Through a systematic factorial study across 96 configurations varying target part-of-speech, retrieval granularity, and sample volume, we identify which factor combinations drive gains and where they break down. Our results demonstrate that static linguistic documentation can be repurposed for machine translation fine-tuning, offering a practical path towards translation tools for severely under-resourced languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。