arXiv:2505.23548cs.CL2025-05被引 4

大模型为何能翻译?研究探讨其背后的数据机制与能力来源。

Translation in the Wild

  • 基于两类预训练数据,模型可能形成双重翻译能力
  • 无需专门训练,仍可实现跨语言理解与转换
  • 适合关注AI翻译原理与语言认知的研究者

大型语言模型(LLMs)在翻译任务中表现出色,尤其在零样本和少样本设置下对多种语言对具备竞争力。然而,与专用神经机器翻译模型不同,这些大模型并未经过任何专门的翻译目标训练。它们为何仍具备强大翻译能力?这种能力是否源于训练数据中的“偶然双语现象”(incidental bilingualism)?指令微调是否发挥了作用?大模型能否识别并利用来自互联网不同角落、语义相同或相似的单语内容,即使这些内容无法同时放入一个上下文窗口?本文结合最新研究与用户经验,提出假设:大模型的翻译能力源自两类不同的预训练数据,可能以不同方式被模型内化。文章讨论了验证这一“二元性”假设的可行性,以及其对深度学习时代翻译概念(人类与机器)的重新思考意义。

原文摘要 · Abstract (English)

Large Language Models (LLMs) excel in translation among other things, demonstrating competitive performance for many language pairs in zero- and few-shot settings. But unlike dedicated neural machine translation models, LLMs are not trained on any translation-related objective. What explains their remarkable translation abilities? Are these abilities grounded in "incidental bilingualism" (Briakou et al. 2023) in training data? Does instruction tuning contribute to it? Are LLMs capable of aligning and leveraging semantically identical or similar monolingual contents from different corners of the internet that are unlikely to fit in a single context window? I offer some reflections on this topic, informed by recent studies and growing user experience. My working hypothesis is that LLMs' translation abilities originate in two different types of pre-training data that may be internalized by the models in different ways. I discuss the prospects for testing the "duality" hypothesis empirically and its implications for reconceptualizing translation, human and machine, in the age of deep learning.

大模型翻译能力语言理解预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。