大模型小模型协作,让高效智能在设备上落地。
A Survey on Collaborative Mechanisms Between Large and Small Language Models
- 通过流水线、路由、蒸馏等机制实现大小模型协同
- 支持低延迟、隐私保护等边缘场景需求
- 适合资源受限环境下的智能应用开发
大语言模型(LLMs)虽具强大能力,但部署成本高、延迟大;小语言模型(SLMs)虽高效易部署,性能有限。大小模型协同成为平衡性能与效率的关键范式,尤其适用于资源受限的边缘设备。本综述系统梳理了LLM-SLM协作的多种交互机制(如流水线、路由、辅助、蒸馏、融合)、关键技术及应用场景,涵盖低延迟、隐私保护、个性化和离线运行等实际需求。尽管展现出构建更高效、可适应、普惠型AI的巨大潜力,仍面临系统开销、模型一致性、任务分配鲁棒性、评估复杂性和安全隐私等挑战。未来方向包括更智能的自适应框架、深度模型融合,以及向多模态和具身智能拓展,推动下一代实用化人工智能发展。
原文摘要 · Abstract (English)
Large Language Models (LLMs) deliver powerful AI capabilities but face deployment challenges due to high resource costs and latency, whereas Small Language Models (SLMs) offer efficiency and deployability at the cost of reduced performance. Collaboration between LLMs and SLMs emerges as a crucial paradigm to synergistically balance these trade-offs, enabling advanced AI applications, especially on resource-constrained edge devices. This survey provides a comprehensive overview of LLM-SLM collaboration, detailing various interaction mechanisms (pipeline, routing, auxiliary, distillation, fusion), key enabling technologies, and diverse application scenarios driven by on-device needs like low latency, privacy, personalization, and offline operation. While highlighting the significant potential for creating more efficient, adaptable, and accessible AI, we also discuss persistent challenges including system overhead, inter-model consistency, robust task allocation, evaluation complexity, and security/privacy concerns. Future directions point towards more intelligent adaptive frameworks, deeper model fusion, and expansion into multimodal and embodied AI, positioning LLM-SLM collaboration as a key driver for the next generation of practical and ubiquitous artificial intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。