arXiv:2605.07244cs.LGcs.AI2026-05

异构大模型通过经验共享实现相互强化,提升训练稳定性与效果。

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models

论文配图:Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models
图 1 · 摘自论文原文
  • 异构大模型在不共享参数的情况下,通过分层设计交换经验。
  • 结果表明,基于结果的共享策略在稳定性和性能间取得最优平衡。
  • 适合研究多模型协作、强化学习后训练的学者与工程师。

我们提出相互强化学习框架,支持异构大语言模型在不共享参数、目标和分词器的前提下,同时进行强化学习后训练。该框架包含共享经验交换(SEE)、多工作资源分配(MWRA)和分词异构层(THL),可对齐不同词汇表间的词级轨迹。在此基础上构建三种受控探测:数据级的同行回放池化(PRP)、价值级的优势共享(XGRPO)和结果级的成功转移(SGT)。上下文带宽分析揭示三者在稳定性-支持性权衡中的位置:PRP承担密度比方差与THL残余成本,XGRPO保持原策略支持但改变标量基线,SGT则提供经验证成功结果的方向引导。在评估设置中,结果级共享位于该权衡的有利位置。

原文摘要 · Abstract (English)

We introduce Mutual Reinforcement Learning, a framework for concurrent RL post-training in which heterogeneous LLM policies exchange typed experience while keeping separate parameters, objectives, and tokenizers. The framework combines a Shared Experience Exchange (SEE), Multi-Worker Resource Allocation (MWRA), and a Tokenizer Heterogeneity Layer (THL) that retokenizes text and aligns token-level traces across incompatible vocabularies. This substrate makes the experience-sharing design question operational across model families. We instantiate three controlled probes on top of GRPO: data-level rollout sharing via Peer Rollout Pooling (PRP), value-level advantage sharing via Cross-Policy GRPO Advantage Sharing (XGRPO), and outcome-level success transfer via Success-Gated Transfer (SGT). A contextual-bandit analysis characterizes their structural positions on a stability-support trade-off: PRP pays density-ratio variance and THL residual costs, XGRPO preserves on-policy actor support while changing scalar baselines, and SGT supplies a rescue-set score direction toward verified peer successes. In the evaluated regime, outcome-level sharing occupies the favorable point of this trade-off.

强化学习异构模型经验共享

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。