arXiv:2502.06567stat.MLcs.LG2025-02被引 3

量化模型会泄露训练数据隐私,该研究提出新方法评估风险。

Membership Inference Risks in Quantized Models: A Theoretical and Empirical Study

  • 基于理论分析设计新指标,衡量量化后模型的隐私暴露程度。
  • 在合成与真实药物数据上验证,可有效比较不同量化器的隐私风险。
  • 适合关注模型隐私安全的研究者与工业界部署人员参考。

量化机器学习模型在降低内存和推理成本的同时,仍能保持与原始模型相当的性能。本文研究量化过程对数据驱动模型隐私的影响,重点关注其对成员推断攻击的脆弱性。会员推断安全(MIS)是近期提出的衡量模型隐私保护能力的新指标,针对最强大(可能未知)的攻击。然而,计算MIS在实际中极为困难。本文提出一种适用于训练后量化流程的新MIS指标,该指标通过最小化经验损失获得。这一指标源自对本情境下MIS的渐近理论分析。我们还提出了一个实证估计该指标的方法。利用合成数据集和真实世界数据(药物发现场景),实验表明该方法能有效评估并排序不同量化器的隐私风险。

原文摘要 · Abstract (English)

Quantizing machine learning models has demonstrated its effectiveness in lowering memory and inference costs while maintaining performance levels comparable to those of the original models. In this work, we investigate the impact of quantization procedures on privacy in data-driven models, focusing on their vulnerability to membership inference attacks. Membership Inference Security (MIS) has recently been proposed to characterize the privacy of machine learning models against the most powerful (and possibly unknown) attacks. However, quantifying MIS appears to be computationally very difficult. In this paper, we propose a new MIS indicator for post-training quantization procedures of machine learning models that minimizes an empirical loss. This new indicator is a byproduct of a theoretical asymptotic analysis of the MIS in this context. We also present a methodology for empirically estimating our MIS indicator. Using synthetic datasets and real-world data (in the context of drug discovery), we demonstrate the effectiveness of our approach in assessing and ranking the MIS of different quantizers.

模型量化隐私安全成员推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。