让大模型与小模型在云端边缘协同训练,提升推理性能。
A Structure-Agnostic Co-Tuning Framework for LLMs and SLMs in Cloud-Edge Systems
- 用代理模型做桥梁,实现大小模型间无结构依赖的知识互学。
- 在跨设备场景下,平均提升 Rouge-L 5.38%、EM 4.88%。
- 适合资源受限的移动端部署,兼顾隐私与实时性需求。
由大语言模型驱动的智能应用迅猛发展,使带宽受限的云服务器难以在不牺牲用户数据隐私的前提下实时处理大规模模型负载。为此,近期研究聚焦于构建融合服务器端大模型与移动边缘设备上小语言模型的云边联合体系。进一步设计此类体系内的协同训练机制以提升推理性能,成为有前景的研究方向。然而,小模型跨域部署及其架构异构性给性能提升带来显著挑战。为此,我们提出 Co-PLMs,一种面向大模型与小模型协同训练的新框架,通过引入无结构依赖的互学机制,实现异构模型间知识交换。该框架利用蒸馏代理模型(DPMs)作为桥梁,促进服务器端 LLM 与设备端 SLM 间的协作训练,同时保留各设备的领域专有信息。实验结果表明,Co-PLMs 超越现有最优方法,在 Rouge-L 上平均提升 5.38%,在 EM 上平均提升 4.88%。
原文摘要 · Abstract (English)
The surge in intelligent applications driven by large language models (LLMs) has made it increasingly difficult for bandwidth-limited cloud servers to process extensive LLM workloads in real time without compromising user data privacy. To solve these problems, recent research has focused on constructing cloud-edge consortia that integrate server-based LLM with small language models (SLMs) on mobile edge devices. Furthermore, designing collaborative training mechanisms within such consortia to enhance inference performance has emerged as a promising research direction. However, the cross-domain deployment of SLMs, coupled with structural heterogeneity in SLMs architectures, poses significant challenges to enhancing model performance. To this end, we propose Co-PLMs, a novel co-tuning framework for collaborative training of large and small language models, which integrates the process of structure-agnostic mutual learning to realize knowledge exchange between the heterogeneous language models. This framework employs distilled proxy models (DPMs) as bridges to enable collaborative training between the heterogeneous server-based LLM and on-device SLMs, while preserving the domain-specific insights of each device. The experimental results show that Co-PLMs outperform state-of-the-art methods, achieving average increases of 5.38% in Rouge-L and 4.88% in EM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。