arXiv:2602.01365cs.LGcs.AI2026-02被引 4

研究不同领域训练顺序对大模型推理能力的影响,发现顺序显著影响效果。

When Domains Interact: Asymmetric and Order-Sensitive Cross-Domain Effects in Reinforcement Learning for Reasoning

  • 按数学、科学等顺序训练,效果远超反向顺序。
  • 先学数学能提升25%准确率,但反向则大幅下降至25%。
  • 没有最优策略,需根据目标领域调整训练顺序。

Group Relative Policy Optimization(GRPO)已成为提升大语言模型推理能力的关键技术,但其在不同领域训练顺序下的表现尚不明确。本文首次系统分析了数学、科学、逻辑和谜题推理任务中多领域训练顺序的影响。结果表明:(1) 单领域泛化具有明显不对称性——先训练其他领域可使数学推理准确率提升约25%,但对逻辑与谜题推理几乎无帮助;(2) 跨领域交互高度依赖顺序——按数学→科学顺序训练时,数学与科学准确率分别为83%和41%,而反向顺序(科学→数学)则分别降至77%和25%;(3) 不存在通用最优策略:顺序训练更利于数学(最高达84%),混合训练更适合科学与逻辑,错误顺序可导致性能差距高达70%到56%。总体表明,多领域环境下GRPO表现出显著的不对称性、顺序敏感性和策略依赖性,强调必须采用领域感知与顺序感知的训练设计。

原文摘要 · Abstract (English)

Group Relative Policy Optimization (GRPO) has become a key technique for improving reasoning abilities in large language models, yet its behavior under different domain sequencing strategies is poorly understood. In particular, the impact of sequential (one domain at a time) versus mixed-domain (multiple domain at a time) training in GRPO has not been systematically studied. We provide the first systematic analysis of training-order effects across math, science, logic, and puzzle reasoning tasks. We found (1) single-domain generalization is highly asymmetric: training on other domains improves math reasoning by approximately 25\% accuracy, while yielding negligible transfer to logic and puzzle; (2) cross-domain interactions are highly order-dependent: training in the order math$\rightarrow$science achieves 83\% / 41\% accuracy on math / science, while reversing the order to science$\rightarrow$math degrades performance to 77\% / 25\%; (3) no single strategy is universally optimal in multi-domain training: sequential training favors math (up to 84\%), mixed training favors science and logic, and poor ordering can incur large performance gaps (from 70\% to 56\%). Overall, our findings demonstrate that GRPO under multi-domain settings exhibits pronounced asymmetry, order sensitivity, and strategy dependence, highlighting the necessity of domain-aware and order-aware training design.

强化学习推理能力训练顺序

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。