用语义奖励强化学习,让大模型学低资源语言不丢通用能力。
Reinforcement Learning with Semantic Rewards Enables Low-Resource Language Expansion without Alignment Tax

- 用嵌入层语义奖励替代传统训练,灵活保留语义而非死记硬背。
- 在藏汉翻译与标题生成中,通用能力遗忘减少52%以上。
- 适合需要安全扩展低资源语言的开发者和研究者。
将大语言模型扩展至低资源语言常导致“对齐代价”:目标语言能力提升却引发通用能力灾难性遗忘。我们指出,这源于监督微调(SFT)的僵化性,其在狭窄且有偏的数据上强制逐词表面模仿。为此,我们提出基于组相对策略优化(GRPO)的语义空间对齐范式,使用嵌入级语义奖励而非似然最大化来优化模型。该目标鼓励通过灵活实现保持语义一致性,实现可控更新,减少对预训练知识的破坏性干扰。我们在藏汉机器翻译和藏文标题生成任务上评估该方法。实验表明,该方法在获得低资源语言能力的同时显著缓解对齐代价,比SFT更有效保留通用能力。尽管表面重合度较低,但语义强化学习在开放生成中产生更高语义质量与偏好,少样本迁移结果也表明其在有限监督下学习到更具迁移性和鲁棒性的表征。总体而言,语义奖励的强化学习为包容性低资源语言扩展提供了更安全、可靠的路径。
原文摘要 · Abstract (English)
Extending large language models (LLMs) to low-resource languages often incurs an "alignment tax": improvements in the target language come at the cost of catastrophic forgetting in general capabilities. We argue that this trade-off arises from the rigidity of supervised fine-tuning (SFT), which enforces token-level surface imitation on narrow and biased data distributions. To address this limitation, we propose a semantic-space alignment paradigm powered by Group Relative Policy Optimization (GRPO), where the model is optimized using embedding-level semantic rewards rather than likelihood maximization. This objective encourages meaning preservation through flexible realizations, enabling controlled updates that reduce destructive interference with pretrained knowledge. We evaluate our approach on Tibetan-Chinese machine translation and Tibetan headline generation. Experiments show that our method acquires low-resource capabilities while markedly mitigating alignment tax, preserving general competence more effectively than SFT. Despite producing less rigid surface overlap, semantic RL yields higher semantic quality and preference in open-ended generation, and few-shot transfer results indicate that it learns more transferable and robust representations under limited supervision. Overall, our study demonstrates that reinforcement learning with semantic rewards provides a safer and more reliable pathway for inclusive low-resource language expansion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。