从统计角度揭示跨语言模型误差根源,降低回答方差可显著提升翻译性能。
Rethinking Cross-lingual Gaps from a Statistical Viewpoint
- 提出跨语言误差由有偏与无偏误差构成,方差是核心成因
- 通过测试时集成减少响应方差,跨语言准确率最高提升12个百分点
- 方法通用性强,适用于多种大模型,适合做多语言知识迁移研究者
网络或大型语料库中的知识通常以一种或少数几种自然语言表达。大语言模型(LLMs)通过从源语言获取知识,并在目标语言查询时提供访问。跨语言差距指在目标语言查询时准确率下降的现象。现有研究聚焦于建模或训练缺陷导致的跨语言差距。本文从统计视角重新审视这一问题,假设目标语言响应的方差是造成差距的关键原因。首次将跨语言差距形式化为有偏误差与无偏误差之和。通过多种推理时干预手段控制方差,实证验证了该假设。提出若干测试时集成方法降低响应方差,使源-目标迁移得分最高提升12个绝对百分点,相对提升达8%至50%以上,涵盖多种LLMs。
原文摘要 · Abstract (English)
Any piece of knowledge is usually expressed in one or a handful of natural languages on the web or in any large corpus. Large Language Models (LLMs) act as a bridge by acquiring knowledge from a source language and making it accessible when queried using target languages. A cross-lingual gap is a drop in accuracy incurred when querying knowledge in a target language rather than the source language. Existing research focused on modeling or training failures leading to cross-lingual gaps. In this work, we take an alternative view to characterize the nature of cross-lingual error, and hypothesize that the variance of responses in the target language is a key cause of this gap. For the first time, we formalize the cross-lingual gap in terms of biased and unbiased errors. We empirically validate our hypothesis through multiple inference-time interventions that control variance and reduce the cross-lingual gap. We demonstrate a few test-time ensemble methods that reduce response variance, and thereby improve source-target transfer scores by up to 12 absolute points yielding relative gains of 8% to over 50% across various LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。