高价值数据在某些攻击中更易被利用,需加强防护。
Understanding Data Importance in Machine Learning Attacks: Does Valuable Data Pose Greater Harm?
- 分析五类攻击,发现重要数据样本更易受侵。
- 高重要性数据在成员推断和模型窃取中漏洞更明显。
- 可结合样本特征提升成员推断效果,适合安全研究者参考。
机器学习在众多领域革新了数据驱动流程,数据对模型性能的影响至关重要。近期研究揭示了个别数据样本的异质性影响,尤其是高价值数据对模型效用的关键贡献。然而,一个关键问题仍未解答:这些高价值数据是否更容易遭受机器学习攻击?本文通过分析五种不同攻击类型,发现高重要性数据样本在某些攻击(如成员推断和模型窃取)中表现出更高的脆弱性。通过分析成员推断脆弱性与数据重要性的关联,研究证明可引入样本特定标准来增强成员推断指标性能。结果凸显了亟需开发兼顾模型效用与高价值数据保护的新防御机制。
原文摘要 · Abstract (English)
Machine learning has revolutionized numerous domains, playing a crucial role in driving advancements and enabling data-centric processes. The significance of data in training models and shaping their performance cannot be overstated. Recent research has highlighted the heterogeneous impact of individual data samples, particularly the presence of valuable data that significantly contributes to the utility and effectiveness of machine learning models. However, a critical question remains unanswered: are these valuable data samples more vulnerable to machine learning attacks? In this work, we investigate the relationship between data importance and machine learning attacks by analyzing five distinct attack types. Our findings reveal notable insights. For example, we observe that high importance data samples exhibit increased vulnerability in certain attacks, such as membership inference and model stealing. By analyzing the linkage between membership inference vulnerability and data importance, we demonstrate that sample characteristics can be integrated into membership metrics by introducing sample-specific criteria, therefore enhancing the membership inference performance. These findings emphasize the urgent need for innovative defense mechanisms that strike a balance between maximizing utility and safeguarding valuable data against potential exploitation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。