研究大模型内部语言一致性对任务表现的影响,发现不一致反而更优。
Do LLMs Need to Think in One Language? Correlation between Latent Language and Task Performance
- 通过多语言提示测试模型内部语言稳定性与任务表现的关系。
- 在翻译和地理文化任务中,内部语言不一致时性能仍稳定甚至更好。
- 适合关注模型内部机制与跨语言能力的研究者阅读。
大型语言模型(LLMs)通常以一种熟练的内部语言(称为隐含语言)处理信息,该语言可能与输入或输出语言不同。然而,隐含语言与输入/输出语言之间的差异如何影响下游任务表现尚不明确。尽管许多研究关注模型的隐含语言,但很少探讨其对任务表现的影响。本文假设:保持隐含语言的一致性有助于提升下游任务表现。为验证此假设,我们在多个下游任务中改变输入提示语言,并分析隐含语言一致性与任务表现之间的相关性。构建了涵盖翻译和地理文化等领域的数据集,这些任务受隐含语言选择的影响较大。在多个LLM上进行的实验结果表明,在翻译和地理文化任务中,维持隐含语言一致性并非最优任务表现的必要条件。这是因为模型在靠近输出层的位置会自适应调整内部表示以匹配目标语言,从而降低了语言一致性对整体性能的影响。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are known to process information using a proficient internal language consistently, referred to as latent language, which may differ from the input or output languages. However, how the discrepancy between the latent language and the input and output language affects downstream task performance remains largely unexplored. While many studies research the latent language of LLMs, few address its importance in influencing task performance. In our study, we hypothesize that thinking in latent language consistently enhances downstream task performance. To validate this, our work varies the input prompt languages across multiple downstream tasks and analyzes the correlation between consistency in latent language and task performance. We create datasets consisting of questions from diverse domains such as translation and geo-culture, which are influenced by the choice of latent language. Experimental results across multiple LLMs on translation and geo-culture tasks, which are sensitive to the choice of language, indicate that maintaining consistency in latent language is not always necessary for optimal downstream task performance. This is because these models adapt their internal representations near the final layers to match the target language, reducing the impact of consistency on overall performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。