arXiv:2412.12500cs.CLcs.AI2024-12中稿 · The First Workshop…被引 11

发现影响多语言模型性能的三大关键因素,助力提升低资源语言表现。

Beyond Data Quantity: Key Factors Driving Performance in Multilingual Language Models

  • 分析204种语言的地理、语言和资源特征,识别核心影响因素。
  • 词元相似性和国家相似性显著提升跨语言迁移效果,优于数据量。
  • 对低资源语言建模有指导意义,适合关注公平性与多语言能力的研究者。

多语言语言模型(MLLMs)在处理多语言文本中至关重要,但因资源分布不均和语言特性差异常出现性能差距。尽管预训练数据比例和模型规模的影响已明确,本研究揭示了其他关键驱动因素。通过分析包括地理、语言和资源在内的多种特征,以SIB-200数据集进行分类、Flores-200数据集进行机器翻译任务,采用回归模型与SHAP值评估204种语言的表现。结果表明,词元相似性与国家相似性是决定模型性能的核心因素,辅以预训练数据量和模型规模。词元相似性促进跨语言知识迁移,而国家相似性凸显共享文化与语言背景的重要性。这些发现为构建更公平、高效的多语言模型提供了重要参考,尤其有助于改善低资源语言的表现。

原文摘要 · Abstract (English)

Multilingual language models (MLLMs) are crucial for handling text across various languages, yet they often show performance disparities due to differences in resource availability and linguistic characteristics. While the impact of pre-train data percentage and model size on performance is well-known, our study reveals additional critical factors that significantly influence MLLM effectiveness. Analyzing a wide range of features, including geographical, linguistic, and resource-related aspects, we focus on the SIB-200 dataset for classification and the Flores-200 dataset for machine translation, using regression models and SHAP values across 204 languages. Our findings identify token similarity and country similarity as pivotal factors, alongside pre-train data and model size, in enhancing model performance. Token similarity facilitates cross-lingual transfer, while country similarity highlights the importance of shared cultural and linguistic contexts. These insights offer valuable guidance for developing more equitable and effective multilingual language models, particularly for underrepresented languages.

多语言模型跨语言迁移低资源语言模型公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。