首个跨领域联邦微调大模型评测基准,助力隐私保护下的领域专用模型开发。
FlowerTune: A Cross-Domain Benchmark for Federated Fine-Tuning of Large Language Models
- 构建跨四大领域的联邦微调评测体系,支持通用NLP、金融、医疗与编程场景。
- 在26个预训练模型上对比不同聚合策略,揭示性能差异与资源约束规律。
- 开源社区驱动,为医疗金融等敏感领域提供可落地的私密微调方案。
大型语言模型(LLMs)在多个领域取得前沿成果,但其发展仍依赖大量公开数据,引发数据稀缺与敏感领域信息难以获取的问题。联邦学习(FL)通过在不共享原始数据的前提下实现去中心化微调,为解决该问题提供了可行框架。然而,预训练LLMs在联邦设置下的兼容性与性能仍缺乏系统研究。本文提出FlowerTune LLM Leaderboard,首个跨领域的联邦微调评测基准,涵盖通用自然语言处理、金融、医疗和编程四个领域。每个领域均配备联邦指令微调数据集及领域专属评估指标。通过协作式开源社区方式,首次对26个预训练模型在不同聚合与微调策略下的联邦表现进行全面比较,揭示模型性能、资源限制与领域适应性之间的关系,为真实场景中隐私保护型、领域专用的大模型研发奠定基础。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have achieved state-of-the-art results across diverse domains, yet their development remains reliant on vast amounts of publicly available data, raising concerns about data scarcity and the lack of access to domain-specific, sensitive information. Federated Learning (FL) presents a compelling framework to address these challenges by enabling decentralized fine-tuning on pre-trained LLMs without sharing raw data. However, the compatibility and performance of pre-trained LLMs in FL settings remain largely under explored. We introduce the FlowerTune LLM Leaderboard, a first-of-its-kind benchmarking suite designed to evaluate federated fine-tuning of LLMs across four diverse domains: general NLP, finance, medical, and coding. Each domain includes federated instruction-tuning datasets and domain-specific evaluation metrics. Our results, obtained through a collaborative, open-source and community-driven approach, provide the first comprehensive comparison across 26 pre-trained LLMs with different aggregation and fine-tuning strategies under federated settings, offering actionable insights into model performance, resource constraints, and domain adaptation. This work lays the foundation for developing privacy-preserving, domain-specialized LLMs for real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。