小模型学会主动求助大模型,动态协作提升推理效率与隐私保护。
Learning to Seek Help: Dynamic Collaboration Between Small and Large Language Models

- 小模型自主决策何时请求大模型协助,实现动态协作。
- 强小模型更独立,强大模型让交互更少但信息量更大。
- 策略可迁移至未见大模型,适合资源受限场景使用。
大型语言模型(LLMs)具备强大能力,但存在成本和隐私问题;小型语言模型(SLMs)支持高效私密的本地推理,但容量有限。为发挥二者互补优势,我们提出一种动态协作框架:在多步推理中,小模型学习主动决定如何向大模型请求帮助,而大模型提供自适应反馈,而非被动工具。我们系统研究了模型能力、效率与隐私约束对协作策略的影响。评估结果揭示显著的缩放效应:更强的小模型更依赖自身,更强的大模型则促进更少但更富信息的交互。所学动态协作策略显著优于静态流水线与独立推理,并能良好迁移至未见过的大模型。
原文摘要 · Abstract (English)
Large language models (LLMs) offer strong capabilities but raise cost and privacy concerns, whereas small language models (SLMs) facilitate efficient and private local inference yet suffer from limited capacity. To synergize the complementary strengths, we introduce a dynamic collaboration framework, where an SLM learns to proactively decide how to request an LLM during multi-step reasoning, while the LLM provides adaptive feedback instead of acting as a passive tool. We further systematically investigate how collaboration strategies are shaped by SLM and LLM capabilities as well as efficiency and privacy constraints. Evaluation results reveal a distinct scaling effect: stronger SLMs become more self-reliant, while stronger LLMs enable fewer and more informative interactions. In addition, the learned dynamic collaboration strategies significantly outperform static pipelines and standalone inference, and transfer robustly to unseen LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。