提出差分隐私语言模型的缩放定律,揭示算力、隐私与性能间权衡关系。
Scaling Laws for Differentially Private Language Models
- 构建差分隐私训练下的模型性能缩放规律
- 明确不同隐私预算下的最优训练配置
- 适合关注隐私保护与模型效率平衡的研究者
缩放定律已成为大语言模型训练中的关键工具,可预测规模扩展带来的性能提升,并指导重要超参数选择,避免高昂试错成本。大语言模型依赖大规模高质量训练数据,如来自用户敏感信息的数据源。在这些敏感数据上训练模型需采用差分隐私(DP)等严格隐私保护机制。然而,差分隐私训练的动力学特性显著不同,其缩放规律尚未完全明晰。本文建立了能准确刻画差分隐私语言模型训练复杂性的缩放定律,全面揭示了算力-隐私-效用之间的权衡关系,并在多种场景下给出了最优训练配置。
原文摘要 · Abstract (English)
Scaling laws have emerged as important components of large language model (LLM) training as they can predict performance gains through scale, and provide guidance on important hyper-parameter choices that would otherwise be expensive. LLMs also rely on large, high-quality training datasets, like those sourced from (sometimes sensitive) user data. Training models on this sensitive user data requires careful privacy protections like differential privacy (DP). However, the dynamics of DP training are significantly different, and consequently their scaling laws are not yet fully understood. In this work, we establish scaling laws that accurately model the intricacies of DP LLM training, providing a complete picture of the compute-privacy-utility tradeoffs and the optimal training configurations in many settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。