用强化学习优化低资源语言模型的分词与翻译质量
VEPO: Variable Entropy Policy Optimization for Low-Resource Language Foundation Models
- 通过可验证奖励机制约束生成结构,确保格式规范
- 在90个语言方向上提升分词效率与翻译质量
- 适合需要提升小语种表现的研究者和开发者
大语言模型在低资源语言上表现不佳,主要源于子词切分效率低和训练数据分布不均。本文提出可变熵策略优化(VEPO),利用带可验证奖励的强化学习,在策略对齐中引入确定性结构约束,确保序列长度固定、格式一致且语言形式正确,全程训练中强制执行。核心是可变熵机制,动态调节字面忠实度与语义自然度的平衡,通过熵加权的优势估计与非对称裁剪,维持有效探索并防止策略坍缩。在FLORES-200、COMET-22及chrF共90个语言方向的实证评估显示,VEPO显著提升分词效率与翻译质量,有效缩小低资源语言的表现差距。
原文摘要 · Abstract (English)
Large language models frequently exhibit suboptimal performance on low resource languages, primarily due to inefficient subword segmentation and systemic training data imbalances. In this paper, we propose Variable Entropy Policy Optimization (VEPO), which leverages Reinforcement Learning with Verifiable Rewards to incorporate deterministic structural constraints into the policy alignment process. This framework ensures prescribed sequence length, robust format consistency, and rigorous linguistic well formedness, all enforced during training. Central to our approach is a variable entropy mechanism that enables the model to dynamically calibrate the equilibrium between literal fidelity and semantic naturalness by modulating the exploration exploitation manifold. By integrating entropy tempered advantage estimation with asymmetric clipping, VEPO sustains robust exploration while mitigating policy collapse. Empirical evaluations across 90 FLORES-200, COMET-22, chrF directions demonstrate that VEPO yields substantial improvements in both tokenization efficiency and translation quality, bridging the performance gap for underrepresented languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。