只需查询模型输出,就能还原训练数据,揭示核方法的隐私风险。
Querying Kernel Methods Suffices for Reconstructing their Training Data
- 通过多次查询模型输出,逆向重构训练数据
- 在核回归、SVM等模型上均验证了重建可行性
- 适合关注机器学习隐私安全的研究者
过参数化模型引发了对记忆训练数据的担忧,即使其泛化性能良好。此类记忆带来的隐私风险在仅能访问模型输出的场景下尚不明确。本文研究核方法中的这一问题,实证与理论证明:仅通过在不同点查询核模型输出,即可重构其训练数据,无需访问模型参数。该结论适用于多种核方法,包括核回归、支持向量机和核密度估计。本工作旨在揭示此类模型潜在的隐私风险。
原文摘要 · Abstract (English)
Over-parameterized models have raised concerns about their potential to memorize training data, even when achieving strong generalization. The privacy implications of such memorization are generally unclear, particularly in scenarios where only model outputs are accessible. We study this question in the context of kernel methods, and demonstrate both empirically and theoretically that querying kernel models at various points suffices to reconstruct their training data, even without access to model parameters. Our results hold for a range of kernel methods, including kernel regression, support vector machines, and kernel density estimation. Our hope is that this work can illuminate potential privacy concerns for such models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。