arXiv:2502.04387cs.CLcs.AI2025-02中稿 · AAAI被引 1

联邦学习自适应优化多语言大模型个性化微调结构

FedP$^2$EFT: Federated Learning to Personalize PEFT for Multilingual LLMs

  • 通过贝叶斯稀疏秩选择,自动为各客户端学习最优的参数高效微调结构
  • 在真实和模拟多语言联邦数据上,性能显著优于现有个性化方法
  • 适合低资源语言场景下的跨设备联邦学习,尤其适用于个性化需求强的场景

联邦学习(FL)已使多语言大语言模型(LLMs)能够在多样且分散的多语言数据上进行训练,特别是在低资源语言场景下。为提升客户端特定性能,常采用参数高效微调(PEFT)模块(如LoRA)进行个性化。这涉及个性化策略(PS),包括PEFT适配器结构设计(例如在哪些层添加LoRA、秩的选择)和超参数选择(如学习率)。现有方法多依赖人工配置,易在低数据环境下过拟合。本文提出FedP²EFT,一种面向跨设备联邦学习场景的多语言大模型个性化微调方法。不同于多数现有结构选择方法,该方法通过贝叶斯稀疏秩选择协同学习每个客户端的最优个性化PEFT结构。在模拟与真实多语言联邦基准上的评估表明,FedP²EFT显著优于现有个性化微调方法,并可与其它联邦学习方法互补。

原文摘要 · Abstract (English)

Federated learning (FL) has enabled the training of multilingual large language models (LLMs) on diverse and decentralized multilingual data, especially on low-resource languages. To improve client-specific performance, personalization via the use of parameter-efficient fine-tuning (PEFT) modules such as LoRA is common. This involves a personalization strategy (PS), such as the design of the PEFT adapter structures (e.g., in which layers to add LoRAs and what ranks) and choice of hyperparameters (e.g., learning rates) for fine-tuning. Instead of manual PS configuration, we propose FedP$^2$EFT, a federated learning-to-personalize method for multilingual LLMs in cross-device FL settings. Unlike most existing PEFT structure selection methods, which are prone to overfitting low-data regimes, FedP$^2$EFT collaboratively learns the optimal personalized PEFT structure for each client via Bayesian sparse rank selection. Evaluations on both simulated and real-world multilingual FL benchmarks demonstrate that FedP$^2$EFT largely outperforms existing personalized fine-tuning methods, while complementing other existing FL methods.

联邦学习多语言参数高效个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。