为模型微调参与提供可量化的个人隐私风险评分
FT-PrivacyScore: Personalized Privacy Scoring Service for Machine Learning Participation
- 基于受控数据访问环境,设计隐私风险量化方法
- 可精准评估个体数据参与微调时的隐私暴露程度
- 适合关注数据隐私的工业与研究场景使用
训练数据隐私是人工智能建模中的首要关切。尽管差异化私密学习等方法允许数据贡献者量化可接受的隐私损失,但常导致模型性能显著下降。实践中,受控数据访问仍是许多工业和研究环境中保护数据隐私的主流方式。在该模式下,授权的模型构建者在受限环境中访问敏感数据,可在降低数据泄露风险的同时完整保留数据效用。然而,与差分隐私不同,当前缺乏对个体数据贡献者在参与机器学习任务前的隐私风险进行定量评估的工具。我们开发了演示原型 FT-PrivacyScore,证明了在模型微调任务中高效、定量估算个人隐私风险的可行性。演示源代码将公开于 https://github.com/RhincodonE/demo_privacy_scoring。
原文摘要 · Abstract (English)
Training data privacy has been a top concern in AI modeling. While methods like differentiated private learning allow data contributors to quantify acceptable privacy loss, model utility is often significantly damaged. In practice, controlled data access remains a mainstream method for protecting data privacy in many industrial and research environments. In controlled data access, authorized model builders work in a restricted environment to access sensitive data, which can fully preserve data utility with reduced risk of data leak. However, unlike differential privacy, there is no quantitative measure for individual data contributors to tell their privacy risk before participating in a machine learning task. We developed the demo prototype FT-PrivacyScore to show that it's possible to efficiently and quantitatively estimate the privacy risk of participating in a model fine-tuning task. The demo source code will be available at \url{https://github.com/RhincodonE/demo_privacy_scoring}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。