提出跨语言迁移理论框架,提升突厥语族低资源语言模型性能
Cross-Lingual Transfer and Parameter-Efficient Adaptation in the Turkic Language Family: A Theoretical Framework for Low-Resource Language Models
- 基于突厥语族语言共性,构建参数高效适配理论模型
- 引入突厥迁移系数量化语言间迁移潜力,最高达0.87
- 适合低资源语言研究者与多语言模型开发者参考
大语言模型虽已革新自然语言处理,但其能力在不同语言间分布不均。多数多语言模型主要基于高资源语言训练,导致大量使用者众多的语言在训练数据和评估基准中仍被忽视,尤其在突厥语族中表现明显。本文提出一种针对突厥语族多语言大模型跨语言迁移与参数高效适配的理论框架,聚焦阿塞拜疆语、哈萨克语、乌兹别克语、土库曼语和加告兹语。这些语言具有高度的形态与句法相似性,但数字资源差异显著,是研究多语言适配策略的理想场景。整合多语言表示学习与参数高效微调技术(如LoRA),构建概念性缩放模型,揭示适配性能如何受模型容量、适配数据量及适配模块表达力影响。为形式化语言间迁移潜力,提出突厥迁移系数(TTC),融合形态相似性、词汇重叠度、句法结构与文字兼容性等维度。框架表明,类型学相似性可促进高效多语言迁移,但也揭示了在极端低资源场景下参数高效适配的结构性局限。
原文摘要 · Abstract (English)
Large language models (LLMs) have transformed natural language processing, yet their capabilities remain uneven across languages. Most multilingual models are trained primarily on high-resource languages, leaving many languages with large speaker populations underrepresented in both training data and evaluation benchmarks. This imbalance is particularly visible in the Turkic language family. This paper proposes a theoretical framework for studying cross-lingual transfer and parameter-efficient adaptation of multilingual LLMs within the Turkic language family, focusing on Azerbaijani, Kazakh, Uzbek, Turkmen, and Gagauz. These languages share substantial typological and morphological similarity while differing greatly in available digital resources, making them a natural setting for analyzing multilingual adaptation strategies. We integrate insights from multilingual representation learning and parameter-efficient fine-tuning techniques such as Low-Rank Adaptation (LoRA) to develop a conceptual scaling model describing how adaptation performance depends on model capacity, adaptation data size, and the expressivity of adaptation modules. To formalize transfer potential between related languages, we introduce the Turkic Transfer Coefficient (TTC), a theoretical measure incorporating morphological similarity, lexical overlap, syntactic structure, and script compatibility across Turkic languages. The framework highlights how typological similarity can enable efficient multilingual transfer while also identifying structural limits of parameter-efficient adaptation in extremely low-resource scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。