用混合奖励与反向迁移调度,让大模型更高效地学多任务。
Omni-Thinker: Scaling Multi-Task RL in LLMs with Hybrid Reward and Task Scheduling
- 混合规则+大模型评优的奖励机制,兼顾确定性与主观任务。
- 多任务性能提升6.2%,比联合训练高,比模型拼接高12.4%。
- 反向迁移预测课程效果,适合做通用大模型后训练的研究者。
通用人工智能的发展依赖于能处理结构化推理与开放生成的大语言模型(LLMs)。我们提出Omni-Thinker,一种统一的强化学习(RL)框架,通过结合混合奖励与基于反向迁移(BWT)的任务调度,实现LLMs在多样化任务上的规模化。混合奖励融合了规则可验证信号与由大模型作为裁判(LLM-as-a-Judge)给出的偏好评估,使模型能在确定性和主观性任务中均有效学习。调度器依据准确率反向迁移(BWT)排序任务,减少遗忘,提升多任务表现。在四个领域的实验表明,该方法相较联合训练提升6.2%,相较模型拼接提升12.4%。此外,我们发现对准确率迁移的简单假设即可精准预测课程结果,而熵动态可解释生成类任务带来的偏差。这些成果凸显了基于BWT的调度与混合监督在推进基于强化学习的后训练以实现通用大模型中的关键作用。
原文摘要 · Abstract (English)
The pursuit of general-purpose artificial intelligence depends on large language models (LLMs) that can handle both structured reasoning and open-ended generation. We present Omni-Thinker, a unified reinforcement learning (RL) framework that scales LLMs across diverse tasks by combining hybrid rewards with backward-transfer-guided scheduling. Hybrid rewards integrate rule-based verifiable signals with preference-based evaluations from an LLM-as-a-Judge, enabling learning in both deterministic and subjective domains. Our scheduler orders tasks according to accuracy backward transfer (BWT), reducing forgetting and improving multi-task performance. Experiments across four domains show gains of 6.2% over joint training and 12.4% over model merging. Moreover, we demonstrate that simple assumptions on accuracy transfer yield accurate predictions of curriculum outcomes, with entropy dynamics explaining deviations due to generative tasks. These findings underscore the importance of BWT-aware scheduling and hybrid supervision for scaling RL-based post-training toward general-purpose LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。