arXiv:2409.20135cs.LGcs.CL2024-09EMNLP被引 1

提升联邦指令微调中跨客户端领域覆盖,显著增强大模型在专业领域的表现。

Optimizing Cross-Client Domain Coverage for Federated Instruction Tuning of Large Language Models

  • 通过选择多样性中心与检索增强,显式最大化跨客户端领域覆盖。
  • 相比11个基线模型,性能最高提升29.19%,领域覆盖提升4.82%~21.36%。
  • 适用于数据稀疏、异构性强的场景,兼顾隐私保护与可扩展性。

针对大语言模型的联邦领域专用指令微调(FedDIT)旨在利用分布式私有且有限的数据提升专业领域表现,但关键性能驱动因素与最优增强策略仍不明确。我们实证发现,跨客户端领域覆盖比数据异质性更具决定性作用。为此提出FedDCA算法,通过面向多样性的客户端中心选择与基于检索的增强,构建多样化、非冗余的跨客户端指令集。在多个领域的大量实验表明,FedDCA优于11个基线模型,性能最高提升29.19%,领域覆盖提升4.82%~21.36%。该方法在数据选择、任务特定公共数据稀缺的保留设置以及不同数据异质性等复杂场景中依然有效,且隐私风险可控。本工作厘清了联邦指令微调的关键动态,提出了一个高效、隐私保护且可扩展的领域专用大模型优化方案。

原文摘要 · Abstract (English)

Federated domain-specific instruction tuning (FedDIT) for large language models (LLMs) aims to enhance performance in specialized domains using distributed private and limited data, yet identifying key performance drivers and optimal augmentation strategies remains challenging. We empirically establish that cross-client domain coverage, rather than data heterogeneity, is the pivotal factor. We then introduce FedDCA, an algorithm that explicitly maximizes this coverage through diversity-oriented client center selection and retrieval-based augmentation, constructing diverse, non-redundant cross-client instruction sets. Extensive experiments across multiple domains demonstrate FedDCA's superiority over eleven baselines, achieving performance gains of up to 29.19\% and domain coverage improvements of 4.82\%-21.36\%. FedDCA maintains its effectiveness in diverse and challenging scenarios, including data selection, held-out settings where task-specific public data is scarce and various data heterogeneity, with manageable privacy risks. This work clarifies critical FedDIT dynamics and presents FedDCA as an effective, privacy-preserving, and scalable solution for advancing domain-specific LLM tuning.

联邦学习指令微调领域适应隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。