arXiv:2602.01928stat.MLcs.LG2026-02

缺失数据能增强隐私保护,首次证明其在差分隐私中的放大作用

Privacy Amplification by Missing Data

  • 将缺失数据视为隐私增强机制,引入差分隐私框架分析
  • 首次证明不完整数据可提升差分隐私算法的隐私保护能力
  • 适用于医疗、金融等敏感数据场景,对数据清洗策略有启示

隐私保护是医疗、金融等高风险领域的重要需求,敏感个人数据需在不泄露个体信息的前提下进行分析。然而,这些应用常因非响应、数据损坏或故意匿名化导致数据缺失。传统观点认为缺失数据是缺陷,会降低信息量并损害模型性能。本文从隐私保护角度重新审视缺失数据:当特征缺失时,个体信息暴露减少,暗示缺失本身可能增强隐私。我们首次在差分隐私框架下形式化这一直觉,证明不完整数据可作为隐私放大机制,为差分隐私算法提供更强的隐私保障。

原文摘要 · Abstract (English)

Privacy preservation is a fundamental requirement in many high-stakes domains such as medicine and finance, where sensitive personal data must be analyzed without compromising individual confidentiality. At the same time, these applications often involve datasets with missing values due to non-response, data corruption, or deliberate anonymization. Missing data is traditionally viewed as a limitation because it reduces the information available to analysts and can degrade model performance. In this work, we take an alternative perspective and study missing data from a privacy preservation standpoint. Intuitively, when features are missing, less information is revealed about individuals, suggesting that missingness could inherently enhance privacy. We formalize this intuition by analyzing missing data as a privacy amplification mechanism within the framework of differential privacy. We show, for the first time, that incomplete data can yield privacy amplification for differentially private algorithms.

隐私保护差分隐私数据缺失

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。