用大模型解决翻译中同义词选择难题,提升译文准确性。
Using Language Models to Disambiguate Lexical Choices in Translation
- 利用大模型生成目标语言的词汇选择规则
- GPT-4在九种语言上准确率达67%至85%
- 规则指导可让弱模型接近甚至超越GPT-4
在翻译中,源语言的一个词可能对应目标语言的多种表达。词汇选择任务需根据上下文判断最合适的变体。我们联合九种语言母语者构建了DTAiLS数据集,包含1,377对句子,展现英译中的跨语言概念差异。评估近期大语言模型与神经机器翻译系统在该数据集上的表现,最佳模型GPT-4在不同语言中准确率为67%至85%。此外,我们使用语言模型生成描述目标语言概念变体的英文规则,将高质量规则提供给性能较弱的模型后,其准确率显著提升,部分情况下达到或超过GPT-4水平。
原文摘要 · Abstract (English)
In translation, a concept represented by a single word in a source language can have multiple variations in a target language. The task of lexical selection requires using context to identify which variation is most appropriate for a source text. We work with native speakers of nine languages to create DTAiLS, a dataset of 1,377 sentence pairs that exhibit cross-lingual concept variation when translating from English. We evaluate recent LLMs and neural machine translation systems on DTAiLS, with the best-performing model, GPT-4, achieving from 67 to 85% accuracy across languages. Finally, we use language models to generate English rules describing target-language concept variations. Providing weaker models with high-quality lexical rules improves accuracy substantially, in some cases reaching or outperforming GPT-4.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。