arXiv:2508.12803cs.CL2025-08被引 2

让语言模型摆脱主流语种干扰,提升方言生成效果

When Alignment Hurts: Decoupling Representational Spaces in Multilingual Models

  • 通过动态探测主流语种子空间,实现生成时的语义解耦
  • 在25种阿拉伯方言上平均提升2.0 chrF++,最高达4.9
  • 适合做多语言生成、低资源语言建模的研究者参考

主流标准语种(如现代标准阿拉伯语)的过度对齐可能损害相关低资源方言的生成能力。本文首次通过因果分析揭示这一现象:大语言模型中主流语种主导的表示子空间会限制方言生成。提出一种在线变分探测框架,在微调过程中持续估计标准语种子空间,并实现投影式解耦。实验基于25种阿拉伯方言的丰富平行数据,结果显示,该方法在保持标准语种性能略有下降的前提下,生成质量平均提升2.0 chrF++,最高达+4.9 chrF++。研究将几何与信息论探测结合,提供可扩展的可控表示调控工具,适用于多语种、跨领域大模型中的表示分配控制。代码将开源。

原文摘要 · Abstract (English)

Alignment with high-resource standard languages is often assumed to aid the modeling of related low-resource varieties. We challenge this assumption by demonstrating that excessive representational entanglement with a dominant variety, such as Modern Standard Arabic (MSA) in relation to Arabic dialects, can actively hinder generative modeling. We present the first comprehensive causal study of this phenomenon by analyzing and directly intervening in the internal representation geometry of large language models (LLMs). Our key contribution is an online variational probing framework that continuously estimates the subspace of the standard variety during fine-tuning, enabling projection-based decoupling from this space. While our study uses Arabic as a case due to its unusually rich parallel resources across 25 dialects, the broader motivation is methodological: dialectal MT serves as a controlled proxy for generative tasks where comparable multi-variety corpora are unavailable. Across 25 dialects, our intervention improves generation quality by up to +4.9 chrF++ and +2.0 on average compared to standard fine-tuning, despite a measured tradeoff in standard-language performance. These results provide causal evidence that subspace dominance by high-resource varieties can restrict generative capacity for related varieties. More generally, we unify geometric and information-theoretic probing with subspace-level causal interventions, offering practical tools for improving generative modeling in closely related language families and, more broadly, for controlling representational allocation in multilingual and multi-domain LLMs. Code will be released.

多语言模型生成质量表示解耦阿拉伯语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。