评估边缘设备上大模型微调的实用可行性,揭示单纯看准确率会误导结论。
EdgeFlowerTune: Evaluating Federated LLM Fine-Tuning Under Realistic Edge System Constraints

- 构建真实设备上的联邦微调基准,综合评估性能与系统开销。
- 发现相同准确率下不同方法在内存、能耗、延迟上差异巨大。
- 适合关注边缘AI部署效率与鲁棒性的研究者和工程师。
联邦微调为在智能手机和物联网设备等边缘端适应大语言模型提供了前景,可在不泄露用户数据隐私的前提下利用丰富的多样化实时数据。这种本地化适配能提升模型个性化、鲁棒性及对本地上下文的响应能力。然而,现有工作多聚焦于跨孤岛或仿真场景,忽视了决定实际可部署性的资源与运行时约束。本文提出EdgeFlowerTune,一个面向真实边缘系统约束的部署导向基准,联合评估模型质量与系统成本(包括通信量、时钟延迟、内存使用、能耗及对动态边缘条件的鲁棒性)。为衡量有效性、效率与鲁棒性,引入三种互补协议:质量-预算、成本-目标与鲁棒性。该基准基于Flower与MobileFineTuner实现,覆盖商用Android手机与NVIDIA边缘开发板。结果表明,仅以准确率为评估标准可能导致错误判断:相似最终质量的方法在实际部署可行性上可能天差地别。EdgeFlowerTune为边缘联邦微调提供可复现的系统感知评估框架。
原文摘要 · Abstract (English)
Federated fine-tuning offers a promising paradigm for adapting large language models (LLMs) on edge devices by leveraging the rich, diverse, and continuously generated data from smartphones and IoT devices without compromising user data privacy. Such edge-side adaptation can improve model personalization, robustness, and responsiveness to local contexts. However, the practical feasibility of federated LLM fine-tuning on real edge devices remains unclear, as most existing work focuses on cross-silo or simulation-based settings, overlooking the resource and runtime constraints that determine whether a method is deployable on real edge systems. We present EdgeFlowerTune, a deployment-oriented benchmark for federated LLM fine-tuning under realistic edge-system constraints. EdgeFlowerTune jointly evaluates model quality and system costs, including communication, wall-clock latency, memory usage, energy consumption, and robustness to dynamic edge conditions. To compare methods in terms of effectiveness, efficiency, and robustness, EdgeFlowerTune introduces three complementary protocols: Quality-under-Budget, Cost-to-Target, and Robustness. We instantiate EdgeFlowerTune as a real-device platform built on Flower and MobileFineTuner, spanning commercial Android smartphones and NVIDIA edge development boards. Our benchmark results show that accuracy-only evaluation can lead to misleading conclusions: methods with similar final quality may differ substantially in deployability once realistic system constraints are considered. EdgeFlowerTune provides a reproducible benchmark for system-aware evaluation of federated LLM fine-tuning at the edge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。