arXiv:2602.09173cs.LGcs.AI2026-02

用强化学习让多个小模型协同推理,效果媲美大模型。

$n$-Musketeers: Reinforcement Learning Shapes Collaboration Among Language Models

  • 通过可训练注意力接口融合多个冻结的小模型内部表示。
  • 在GSM8K上表现接近单模型强化学习基线,复杂任务中专家注意力更集中。
  • 揭示了不同任务下专家协作模式的演变,适合研究模型协同机制的人。

近期基于可验证奖励的强化学习(RLVR)进展表明,小型专用语言模型(SLMs)可在不依赖大型单一模型的情况下展现出结构化推理能力。我们提出软隐藏状态协作机制,通过可训练的注意力接口,将多个异构的冻结SLM专家基于其内部表征进行集成。在Reasoning Gym和GSM8K上的实验显示,这种隐式集成方法在性能上可与强大的单模型RLVR基线相媲美。消融实验进一步揭示了专家利用的双重机制:在简单算术任务中,性能提升主要由静态专家偏好解释;而在更复杂任务中,随着训练推进,专家注意力逐渐集中并形成结构性分配,表明路由机制能催生专家的涌现专业化。总体而言,隐藏状态协作提供了一种紧凑的冻结专家利用方式,并为观察专家使用模式及其在RLVR下的演化提供了窗口。

原文摘要 · Abstract (English)

Recent progress in reinforcement learning with verifiable rewards (RLVR) shows that small, specialized language models (SLMs) can exhibit structured reasoning without relying on large monolithic LLMs. We introduce soft hidden-state collaboration, where multiple heterogeneous frozen SLM experts are integrated through their internal representations via a trainable attention interface. Experiments on Reasoning Gym and GSM8K show that this latent integration is competitive with strong single-model RLVR baselines. Ablations further reveal a dual mechanism of expert utilization: for simpler arithmetic domains, performance gains can largely be explained by static expert preferences, whereas more challenging settings induce increasingly concentrated and structured expert attention over training, indicating emergent specialization in how the router connects to relevant experts. Overall, hidden-state collaboration provides a compact mechanism for leveraging frozen experts, while offering an observational window into expert utilization patterns and their evolution under RLVR.

强化学习模型协同小模型推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。