综述大模型在上下文感知机器翻译中的应用与前景
Beyond the Sentence: A Survey on Context-Aware Machine Translation with Large Language Models
- 梳理大模型在上下文翻译中的提示与微调方法
- 商业模型如ChatGPT优于开源模型如Llama
- 适合关注LLM翻译性能与未来方向的研究者
尽管大语言模型(LLMs)广受欢迎,其在机器翻译领域的应用仍相对不足,尤其在上下文感知场景下。本文对基于大模型的上下文感知翻译研究进行综述。现有工作主要采用提示工程和微调方法,少数聚焦自动后编辑与翻译代理构建。我们发现,商用大模型(如ChatGPT、Tower LLM)表现优于开源模型(如Llama、Bloom LLM),且提示方法可作为翻译质量评估的良好基线。最后,文章提出若干值得探索的未来方向。
原文摘要 · Abstract (English)
Despite the popularity of the large language models (LLMs), their application to machine translation is relatively underexplored, especially in context-aware settings. This work presents a literature review of context-aware translation with LLMs. The existing works utilise prompting and fine-tuning approaches, with few focusing on automatic post-editing and creating translation agents for context-aware machine translation. We observed that the commercial LLMs (such as ChatGPT and Tower LLM) achieved better results than the open-source LLMs (such as Llama and Bloom LLMs), and prompt-based approaches serve as good baselines to assess the quality of translations. Finally, we present some interesting future directions to explore.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。