arXiv:2608.23358cs.CL2026-08

低资源语言在大模型中表征退化,几何正则化可有效改善。

The Geometry of Low-Resource Language Representations

  • 从表征几何角度分析模型内部差异,发现低资源语言最终层表征退化明显。
  • 在9个基模型适配10种非洲语言的实验中,几何正则化显著减少退化现象。
  • 对困难任务效果更优,适合低资源语言持续预训练场景。

大模型在低资源与高资源语言间的性能差距广为人知,但其内在驱动因素仍不明确。本文通过表征几何视角刻画这一差距。对比30种语言的隐藏表征几何特性发现,模型几何结构与语言数据量系统相关。最显著的影响出现在最终层,低资源语言表现出表征退化。为缓解此问题,我们研究了在持续预训练(CPT)中引入正则化项以惩罚退化的有效性。在9个基础模型单语微调10种非洲语言的实验中,几何正则化成功减少了表征退化。对于大模型,基于余弦相似性的正则化在性能上略优于原始CPT,尤其在最具挑战性的任务上表现更稳定。结果表明,低资源与高资源语言在大模型中的表征几何存在可测量差异,且针对性的几何干预是提升低资源语言持续预训练的有效策略。

原文摘要 · Abstract (English)

The performance gap between low- and high-resource languages in LLMs is widely known, but it remains unclear which internal model factors drive these disparities. In this paper, we characterise this gap through the lens of representational geometry. Comparing the geometric properties of hidden representations across 30 languages reveals that LLM geometry is systematically related to language data availability. The most consistent effect is in final layers, where low-resource languages exhibit representational degeneration. To counter this, we investigate the effectiveness of regularisation terms to penalise degeneration during continued pretraining (CPT). Experiments monolingually adapting 9 base LLMs to 10 African languages show that geometric regularisation successfully reduces representational degeneration during CPT. For larger models, cosine similarity-based regularisation marginally improves performance over vanilla CPT, with more consistent gains on the most challenging tasks. We establish that the representational geometry of low- and high-resource languages in LLMs is measurably distinct, and that targeted geometric intervention is a viable strategy for improving CPT for low-resource languages.

表征几何低资源语言持续预训练正则化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。