arXiv:2512.08777cs.CLcs.AI2025-12

让低资源语言模型在不流畅评分器下仍保持语言流畅性。

Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages

  • 用在线策略训练法,避免依赖目标语种的指令数据。
  • 挪威语实验中,流畅度优于翻译微调和多语言微调。
  • 适合缺乏本族语数据的低资源语言模型优化。

我们提出一种面向低资源语言的后训练方法,使语言模型在由不流畅奖励模型对齐时仍能保持语言流畅性。偏好优化已成研究热点,但以往工作主要集中于英语和中文。低资源语言既缺乏母语者撰写的语料,也缺少能生成流畅合成数据的指令微调模型。为此,我们致力于构建无需目标语言指令微调数据的流畅偏好对齐语言模型。方法采用在线策略训练,与两种常见替代方案——基于机器翻译的监督微调和多语言微调——进行对比。以挪威语博克马尔文为案例,通过母语者评估验证流畅度。结果表明,在线策略至关重要,且无需依赖难以获取的数据即可超越其他方法。

原文摘要 · Abstract (English)

We propose a post-training method for lower-resource languages that preserves the fluency of language models even when aligned by disfluent reward models. Preference optimization is now a well-researched topic, but previous work has mostly addressed models for English and Chinese. Lower-resource languages lack both datasets written by native speakers and instruction-tuned language models capable of generating fluent synthetic data. To address this, we focus on developing a fluent preference-aligned language model without any instruction-tuning data in the target language. Our approach uses an on-policy training method, which we compare with two common alternatives: supervised finetuning on machine-translated data and multilingual finetuning. We conduct a case study on Norwegian Bokmål and evaluate fluency through native-speaker assessments. The results show that the on-policy aspect is crucial and outperforms the alternatives without relying on any hard-to-obtain data.

低资源语言偏好对齐语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。