对比多语言大模型表现,发现资源少的语言能力明显更弱。
Exploring Multilingual Probing in Large Language Models: A Cross-Language Analysis
- 跨语言探针分析不同语言下模型表现差异
- 高资源语言探针准确率显著高于低资源语言
- 低资源语言间及与高资源语言间表征相似性较低
大语言模型的探针技术长期集中于英语,忽略了世界大多数语言。本文将探针方法拓展至多语言场景,对多个开源大模型进行实验,分析探针准确率、层间趋势以及多语言探针向量的相似性。关键发现:(1) 高资源语言与低资源语言之间存在持续性能差距,前者探针准确率显著更高;(2) 层级趋势分化,高资源语言在深层中表现提升明显,类似英语;(3) 高资源语言间表征相似性较高,而低资源语言彼此间及与高资源语言间相似性均较低。结果揭示了大模型在多语言能力上的显著不平等,强调需改进对低资源语言的建模。
原文摘要 · Abstract (English)
Probing techniques for large language models (LLMs) have primarily focused on English, overlooking the vast majority of the world's languages. In this paper, we extend these probing methods to a multilingual context, investigating the behaviors of LLMs across diverse languages. We conduct experiments on several open-source LLM models, analyzing probing accuracy, trends across layers, and similarities between probing vectors for multiple languages. Our key findings reveal: (1) a consistent performance gap between high-resource and low-resource languages, with high-resource languages achieving significantly higher probing accuracy; (2) divergent layer-wise accuracy trends, where high-resource languages show substantial improvement in deeper layers similar to English; and (3) higher representational similarities among high-resource languages, with low-resource languages demonstrating lower similarities both among themselves and with high-resource languages. These results highlight significant disparities in LLMs' multilingual capabilities and emphasize the need for improved modeling of low-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。