通过层间对齐提升低资源语言翻译质量
TRepLiNa: Layer-wise CKA+REPINA Alignment Improves Low-Resource Machine Translation in Aya-23 8B
- 结合CKA与REPINA,在解码器中层强制跨语言表示对齐
- 在数据稀缺场景下,对曼达里、桑塔利等语言翻译效果显著提升
- 适合低资源语言研究者,且无需大量标注数据
2025年多模态模型在低资源语境与社会影响(MMLoSo)语言挑战赛聚焦印度多样低资源语言(LRLs)缺乏资源的核心问题。本研究探索在仅解码器的多语言大模型特定内部层中强制跨语言相似性,是否能提升从低资源语言到高资源语言(HRL)的翻译质量。我们提出TRepLiNa方法,将中心核对齐(CKA)与参数更新正则化方法REPINA相结合。实验基于Aya-23 8B模型,采用QLoRA,在MMLoSo共享任务的语言对(曼达里、桑塔利、比利语)与印地语/英语互译场景下,测试零样本、少样本及微调设置。结果表明,使用TRepLiNa对中间层进行对齐是一种低成本、实用的方法,在数据稀缺情况下显著改善低资源语言翻译性能。
原文摘要 · Abstract (English)
The 2025 Multimodal Models for Low-Resource Contexts and Social Impact (MMLoSo) Language Challenge addresses one of India's most pressing linguistic gaps: the lack of resources for its diverse low-resource languages (LRLs). In this study, we investigate whether enforcing cross-lingual similarity in specific internal layers of a decoder-only multilingual large language model (LLM) can improve translation quality from LRL to high-resource language (HRL). Specifically, we combine Centered Kernel Alignment (CKA), a similarity metric that encourages representations of different languages to align, with REPINA, a regularization method that constrains parameter updates to remain close to the pretrained model, into a joint method we call TRepLiNa. In this research project, we experiment with zero-shot, few-shot, and fine-tuning settings using Aya-23 8B with QLoRA across MMLoSo shared task language pairs (Mundari, Santali, Bhili) with Hindi/English pivots. Our results show that aligning mid-level layers using TRepLiNa (CKA+REPINA) is a low-cost, practical approach to improving LRL translation, especially in data-scarce settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。