arXiv:2602.14302cs.DCcs.LG2026-02中稿 · IEEE Transactions …被引 5

Floe让边缘设备用轻量模型实时推理,隐私数据留在本地。

Floe: Federated Specialization for Real-Time LLM-SLM Inference

  • 云上大模型+边缘小模型协同,本地不传数据
  • 多硬件适配高效微调,边缘延迟降低40%以上
  • 适合对隐私和响应速度要求高的智能终端

在资源受限、延迟敏感的环境中部署大语言模型(LLMs)仍具挑战性,主要源于其巨大的计算开销和隐私风险。本文提出Floe,一种混合联邦学习框架,通过云端黑箱式大模型与边缘设备上的轻量级小语言模型(SLMs)协同,实现低延迟、隐私保护的推理。用户私有数据及微调过程均保留在设备端,云端模型仅提供通用知识而不暴露自有权重。采用异构感知的LoRA适配策略,可高效支持多种硬件配置;同时引入基于输出概率的融合机制,实现边缘与云端模型的实时协同。大量实验表明,Floe在保障用户隐私与个性化的同时,显著提升边缘设备的模型性能,并在实时约束下将推理延迟降低40%以上,优于基线方法。

原文摘要 · Abstract (English)

Deploying large language models (LLMs) in real-time systems remains challenging due to their substantial computational demands and privacy concerns. We propose Floe, a hybrid federated learning framework designed for latency-sensitive, resource-constrained environments. Floe combines a cloud-based black-box LLM with lightweight small language models (SLMs) on edge devices to enable low-latency, privacy-preserving inference. Personal data and fine-tuning remain on-device, while the cloud LLM contributes general knowledge without exposing proprietary weights. A heterogeneity-aware LoRA adaptation strategy enables efficient edge deployment across diverse hardware, and a logit-level fusion mechanism enables real-time coordination between edge and cloud models. Extensive experiments demonstrate that Floe enhances user privacy and personalization. Moreover, it significantly improves model performance and reduces inference latency on edge devices under real-time constraints compared with baseline approaches.

联邦学习边缘推理隐私保护小模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。