测试大模型在文档翻译中是否真能用好上下文。
Context-Aware or Context-Insensitive? Assessing LLMs' Performance in Document-Level Translation
- 通过扰动和归因分析,评估大模型对文档上下文的依赖程度。
- 大模型整体表现优于传统模型,但代词翻译仍不靠谱。
- 需针对性微调以提升模型对关键上下文的感知能力,适合文档翻译研究者参考。
大型语言模型(LLMs)在机器翻译中表现日益突出。本文聚焦文档级翻译,研究主流大模型在翻译时利用文档上下文的能力。通过扰动分析(考察模型对被扰动或随机化文档上下文的鲁棒性)和归因分析(检验相关上下文对翻译结果的贡献),我们在九个来自不同模型家族和训练范式的LLM上进行了广泛评估,包括专门用于翻译的LLM以及两个编码器-解码器Transformer基线模型。结果表明,尽管LLMs在文档翻译任务上的性能优于编码器-解码器模型,但其在代词翻译方面的表现并未体现出相应优势。分析揭示了需要针对文档上下文中相关部分进行上下文感知的微调,以提高大模型在文档级翻译中的可靠性。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly strong contenders in machine translation. In this work, we focus on document-level translation, where some words cannot be translated without context from outside the sentence. Specifically, we investigate the ability of prominent LLMs to utilize the document context during translation through a perturbation analysis (analyzing models' robustness to perturbed and randomized document context) and an attribution analysis (examining the contribution of relevant context to the translation). We conduct an extensive evaluation across nine LLMs from diverse model families and training paradigms, including translation-specialized LLMs, alongside two encoder-decoder transformer baselines. We find that LLMs' improved document-translation performance compared to encoder-decoder models is not reflected in pronoun translation performance. Our analysis highlight the need for context-aware finetuning of LLMs with a focus on relevant parts of the context to improve their reliability for document-level translation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。