arXiv:2410.15956cs.CLcs.AI2024-10ACL被引 40

发现大模型在非英语中常有‘英语口音’,提出新方法让多语言输出更自然。

Do Large Language Models Have an English Accent? Evaluating and Improving the Naturalness of Multilingual LLMs

  • 设计跨语言自然度评估指标,检测非英语输出的语法词汇异常
  • 在法语和中文数据集上发现主流模型存在英语模式偏差
  • 提出对齐方法提升目标语言自然度,且不损害通用能力

当前大型语言模型主要以英语为核心设计,即使少数多语言模型也普遍存在英语中心偏差。如同二语使用者常出现表达生硬的情况,这些模型在非英语语言中生成的文本往往表现出词汇与语法层面的不自然,反映出强烈的英语影响。尽管该问题至关重要,但现有研究对其关注有限。本文引入新的自动语料级评估指标,从词汇和句法层面量化多语言输出的自然度。基于这些指标,在法语与中文的精选基准上评估主流大模型,发现其普遍呈现英语主导特征。为缓解此问题,我们提出一种简单有效的对齐方法,显著提升模型在目标语言与领域中的自然度,且不影响其在通用基准上的表现。本工作强调了构建多语言评估体系与优化方法的重要性,推动下一代多语言大模型的发展。

原文摘要 · Abstract (English)

Current Large Language Models (LLMs) are predominantly designed with English as the primary language, and even the few that are multilingual tend to exhibit strong English-centric biases. Much like speakers who might produce awkward expressions when learning a second language, LLMs often generate unnatural outputs in non-English languages, reflecting English-centric patterns in both vocabulary and grammar. Despite the importance of this issue, the naturalness of multilingual LLM outputs has received limited attention. In this paper, we address this gap by introducing novel automatic corpus-level metrics to assess the lexical and syntactic naturalness of LLM outputs in a multilingual context. Using our new metrics, we evaluate state-of-the-art LLMs on a curated benchmark in French and Chinese, revealing a tendency towards English-influenced patterns. To mitigate this issue, we also propose a simple and effective alignment method to improve the naturalness of an LLM in a target language and domain, achieving consistent improvements in naturalness without compromising the performance on general-purpose benchmarks. Our work highlights the importance of developing multilingual metrics, resources and methods for the new wave of multilingual LLMs.

多语言模型自然度评估语言偏见模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。