用语法标注提升机器翻译,零训练成本,低资源语言效果显著
GrammaMT: Improving Machine Translation with Grammar-Informed In-Context Learning
- 用语义标注文本引导大模型翻译,不需训练
- 在濒危语言任务上提升超17个BLEU点
- 适合低资源语言和小样本场景
我们提出GrammaMT,一种基于语法标注的提示方法,利用跨行词典文本(IGT)对源句进行形态和词汇注释。GrammaMT设计了三种无需训练的提示策略:gloss-shot、chain-gloss 和 model-gloss,仅需少量人工标注样本即可部署,适用于低资源场景。实验表明,该方法在三个基准测试中均提升了开源指令微调大模型的翻译性能:(1)最大规模的IGT语料库;(2)2023 SIGMORPHON共享任务中濒危语言数据集;(3)以及在FLORES外域设置下的表现。消融实验显示,若大模型能准确生成或访问源句的标注信息,翻译性能可提升超过17 BLEU点。
原文摘要 · Abstract (English)
We introduce GrammaMT, a grammatically-aware prompting approach for machine translation that uses Interlinear Glossed Text (IGT), a common form of linguistic description providing morphological and lexical annotations for source sentences. GrammaMT proposes three prompting strategies: gloss-shot, chain-gloss and model-gloss. All are training-free, requiring only a few examples that involve minimal effort to collect, and making them well-suited for low-resource setups. Experiments show that GrammaMT enhances translation performance on open-source instruction-tuned LLMs for various low- to high-resource languages across three benchmarks: (1) the largest IGT corpus, (2) the challenging 2023 SIGMORPHON Shared Task data over endangered languages, and (3) even in an out-of-domain setting with FLORES. Moreover, ablation studies reveal that leveraging gloss resources could substantially boost MT performance (by over 17 BLEU points) if LLMs accurately generate or access input sentence glosses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。