arXiv:2602.00166cs.LGcs.AI2026-02被引 2

让小模型学会在预算内智能调用大模型,避免遗忘且行为稳定。

Joint Continual Learning of Local Language Models and Cloud Offloading Decisions with Budget Constraints

  • 通过双优势机制动态调整云调用决策,不依赖固定奖励设计。
  • 在数学推理与代码生成任务上,准确率提升且遗忘显著减少。
  • 适合资源受限场景下持续学习的小模型协同部署应用。

本地部署的小语言模型(SLMs)需在严格内存与计算约束下持续支持多样化任务,因此选择性依赖云端大语言模型(LLMs)不可避免。在持续学习过程中,调控云端协助极具挑战,因朴素的基于奖励的强化学习常导致调用行为不稳定,并加剧任务分布变化时的灾难性遗忘。我们提出DA-GRPO,即组相对策略优化的双优势扩展,将云使用约束直接融入优势计算,避免固定奖励塑造与外部路由模型。该设计使本地模型能联合学习任务能力与协作行为,使云端请求在后训练阶段自然涌现,同时遵守预设协助预算。在数学推理与代码生成基准上的实验表明,相较于先前的协作与路由方法,DA-GRPO提升了切换后的准确率,大幅降低遗忘,且保持稳定的云使用水平。

原文摘要 · Abstract (English)

Locally deployed Small Language Models (SLMs) must continually support diverse tasks under strict memory and computation constraints, making selective reliance on cloud Large Language Models (LLMs) unavoidable. Regulating cloud assistance during continual learning is challenging, as naive reward-based reinforcement learning often yields unstable offloading behavior and exacerbates catastrophic forgetting as task distributions shift. We propose DA-GRPO, a dual-advantage extension of Group Relative Policy Optimization that incorporates cloud-usage constraints directly into advantage computation, avoiding fixed reward shaping and external routing models. This design enables the local model to jointly learn task competence and collaboration behavior, allowing cloud requests to emerge naturally during post-training while respecting a prescribed assistance budget. Experiments on mathematical reasoning and code generation benchmarks show that DA-GRPO improves post-switch accuracy, substantially reduces forgetting, and maintains stable cloud usage compared to prior collaborative and routing-based approaches.

持续学习云协同小模型预算约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。