用特权对比判别器提升多语言推理,无需目标语言数据
Gained in Translation: Privileged Pairwise Judges Enhance Multilingual Reasoning
- 两阶段训练:先翻译微调,再用带参考答案的对比反馈强化
- 在数学和非数学任务上超越全量训练模型,仅需少量数据
- 适合缺乏目标语言数据的多语言推理研究者
当面对训练数据中少见的语言时,当前推理大模型的表现会显著下降。为此,我们提出 exttt{SP3F}(自对弈带特权对比反馈),一种无需目标语言数据的两阶段多语言推理增强框架。首先,在英文问答对的翻译版本上进行监督微调以提升基础模型正确率;其次,采用自对弈强化学习,判别器在获得英文参考答案作为特权信息的前提下,判断模型输出的优劣,即使两者均不完全正确也能选出更优者。整体上, exttt{SP3F} 显著提升基线模型性能,在单语言、多语言及未见语言泛化设置下,均优于全量后训练模型,且训练数据用量不足其十分之一。
原文摘要 · Abstract (English)
When asked a question in a language less seen in its training data, current reasoning large language models (RLMs) often exhibit dramatically lower performance than when asked the same question in English. In response, we introduce \texttt{SP3F} (Self-Play with Privileged Pairwise Feedback), a two-stage framework for enhancing multilingual reasoning without \textit{any} data in the target language(s). First, we supervise fine-tune (SFT) on translated versions of English question-answer pairs to raise base model correctness. Second, we perform RL with feedback from a pairwise judge in a self-play fashion, with the judge receiving the English reference response as \textit{privileged information}. Thus, even when none of the model's responses are completely correct, the privileged pairwise judge can still tell which response is better. End-to-end, \texttt{SP3F} greatly improves base model performance, even outperforming fully post-trained models on multiple math and non-math tasks with less than of the training data across the single-language, multilingual, and generalization to unseen language settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。