联邦学习中应优先使用开源模型进行语言模型后训练
Federated Foundation Language Model Post-Training Should Focus on Open-Source Models
- 主张联邦后训练采用开源模型而非黑盒模型
- 黑盒模型违背数据隐私与自主性核心原则
- 适合关注隐私保护与模型可控性的研究者
基础语言模型的后训练已成为联邦学习中一个有前景的研究方向,旨在实现隐私保护下的模型优化与用户下游任务适配。近期进展多采用基于黑盒基础语言模型的集中式后训练方法,无法访问模型权重和架构细节。尽管黑盒模型在集中式后训练中表现良好,但在联邦学习中直接复用引发诸多担忧。我们认为,联邦学习中使用黑盒模型违背了数据隐私与自主性的核心原则。本文批判性分析了黑盒模型在联邦后训练中的应用,详细阐述了开放性的多个方面及其对联邦学习的影响。
原文摘要 · Abstract (English)
Post-training of foundation language models has emerged as a promising research domain in federated learning (FL) with the goal to enable privacy-preserving model improvements and adaptations to user's downstream tasks. Recent advances in this area adopt centralized post-training approaches that build upon black-box foundation language models where there is no access to model weights and architecture details. Although the use of black-box models has been successful in centralized post-training, their blind replication in FL raises several concerns. Our opinion is that using black-box models in FL contradicts the core principles of federation such as data privacy and autonomy. In this paper, we critically analyze the usage of black-box models in federated post-training, and provide a detailed account of various aspects of openness and their implications for FL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。