arXiv:2512.03976cs.CL2025-12被引 3

用两阶段方法让大模型学会藏语,翻译质量提升显著。

Adapting Large Language Models to Low-Resource Tibetan: A Two-Stage Continual and Supervised Fine-Tuning Study

  • 先持续预训练建语言基础,再监督微调做翻译专精。
  • 困惑度从2.98降到1.54,翻译BLEU值提升至0.261。
  • 适合研究低资源语言适配或想复现藏语模型的人。

将大语言模型(LLM)适配至低资源语言仍面临数据稀缺与跨语言漂移的挑战。本文针对语法复杂、资源匮乏的藏语,对Qwen2.5-3B模型进行两阶段适配:首先通过持续预训练(CPT)建立藏语语言基础,随后通过监督微调(SFT)实现任务与翻译专业化。实验表明,困惑度由2.98降至1.54,中译藏翻译质量显著提升(BLEU:0.046→0.261;chrF:2.2→6.6)。对Qwen3-4B模型435层的分层分析显示,适应主要集中在嵌入层和输出头,中后期MLP投影编码领域特异性变换。结果表明,CPT构建了藏语语义空间,而SFT在最小扰动下强化任务对齐。本研究首次定量揭示藏语适配动态,提供可复现的多语言大模型扩展框架。

原文摘要 · Abstract (English)

Adapting large language models (LLMs) to low-resource languages remains a major challenge due to data scarcity and cross-lingual drift. This work presents a two-stage adaptation of Qwen2.5-3B to Tibetan, a morphologically rich and underrepresented language. We employ Continual Pretraining (CPT) to establish Tibetan linguistic grounding, followed by Supervised Fine-Tuning (SFT) for task and translation specialization. Empirical evaluations demonstrate a consistent decrease in perplexity (from 2.98 $\rightarrow$ 1.54) and substantial improvements in Chinese$\rightarrow$Tibetan translation quality (BLEU: 0.046 $\rightarrow$ 0.261; chrF: 2.2 $\rightarrow$ 6.6). Layer-wise analysis across 435 layers in Qwen3-4B reveals that adaptation primarily concentrates on embedding and output heads, with mid--late MLP projections encoding domain-specific transformations. Our findings suggest that CPT constructs a Tibetan semantic manifold while SFT sharpens task alignment with minimal representational disruption. This study provides the first quantitative exploration of Tibetan adaptation dynamics for LLMs, and offers an open, reproducible framework for extending multilingual foundation models to low-resource settings.

大模型藏语低资源微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。