分析图像数据集特征如何影响隐私保护模型的性能与安全。
Assessing the Impact of Image Dataset Features on Privacy-Preserving Machine Learning
- 研究不同数据集特性对隐私模型的影响,发现不平衡数据会放大少数类风险。
- 低类别数数据提升模型效用与隐私保护,高熵或低FDR数据则恶化平衡性。
- 为数据选择与隐私优化提供可操作指导,适合关注隐私安全的研究者。
机器学习在计算机视觉等领域至关重要,但训练于敏感数据的模型面临安全威胁,可能被攻击导致信息泄露。隐私保护机器学习(PPML)通过差分隐私(DP)在模型效用与隐私间寻求平衡。本研究分析了多个数据集及不同隐私预算下,图像数据集特征对私有与非私有卷积神经网络(CNN)模型效用与脆弱性的影响。结果表明:数据分布不均会加剧少数类的脆弱性,但差分隐私可缓解此问题;类别较少的数据集能同时提升模型效用与隐私保护能力;而高熵或低费雪判别比(FDR)的数据集则恶化效用-隐私权衡。这些发现为从业者和研究人员评估与优化图像数据集中的效用-隐私平衡提供了重要依据,有助于基于数据特征进行数据或隐私策略调整以获得更优结果。
原文摘要 · Abstract (English)
Machine Learning (ML) is crucial in many sectors, including computer vision. However, ML models trained on sensitive data face security challenges, as they can be attacked and leak information. Privacy-Preserving Machine Learning (PPML) addresses this by using Differential Privacy (DP) to balance utility and privacy. This study identifies image dataset characteristics that affect the utility and vulnerability of private and non-private Convolutional Neural Network (CNN) models. Through analyzing multiple datasets and privacy budgets, we find that imbalanced datasets increase vulnerability in minority classes, but DP mitigates this issue. Datasets with fewer classes improve both model utility and privacy, while high entropy or low Fisher Discriminant Ratio (FDR) datasets deteriorate the utility-privacy trade-off. These insights offer valuable guidance for practitioners and researchers in estimating and optimizing the utility-privacy trade-off in image datasets, helping to inform data and privacy modifications for better outcomes based on dataset characteristics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。