用语法结构提升古科普特语翻译效果,结合词典与句法信息实现新纪录。
Syntax as a Rosetta Stone: Universal Dependencies for In-Context Coptic Translation

- 输入中加入通用依存句法分析结果,增强模型理解能力。
- 词典+句法信息组合在不同规模模型上均显著提升翻译性能。
- 适合低资源语言翻译研究者,尤其关注古语言处理场景。
低资源机器翻译需采用区别于高资源语言的方法。本文提出一种新颖的上下文学习方法,用于支持古科普特语到英语的低资源机器翻译,通过通用依存句法(Universal Dependencies)对输入句子进行句法增强。在现有基于双语词典支持词汇项推理的基础上,我们在输入中引入多种句法表示:原始解析输出、用自然英语描述的句法结构,以及针对子树中难点结构的目标指令。实验表明,尽管仅使用句法信息不如词典标注有效,但将检索到的词典项与句法信息结合后,在不同模型规模下均取得显著提升,创下古科普特语翻译的新基准表现。
原文摘要 · Abstract (English)
Low-resource machine translation requires methods that differ from those used for high-resource languages. This paper proposes a novel in-context learning approach to support low-resource machine translation of the Coptic language to English, with syntactic augmentation from Universal Dependencies parses of input sentences. Building on existing work using bilingual dictionaries to support inference for vocabulary items, we add several representations of syntactic analyses to our inputs , specifically exploring the inclusion of raw parser outputs, verbalizations of parses in plain English, and targeted instructions of difficult constructions identified in sub-trees and how they can be translated. Our results show that while syntactic information alone is not as useful as dictionary-based glosses, combining retrieved dictionary items with syntactic information achieves significant gains across model sizes, achieving new state-of-the-art translation results for Coptic.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。