不切分数据也能做隐私保护的置信预测,效果更优。
Beyond Data Splitting: Full-Data Conformal Prediction by Differential Privacy
- 用差分隐私保证稳定性,避免数据切分
- 实验显示预测集比传统方法更紧凑
- 适合需要高精度与隐私保护的场景
隐私保护与不确定性量化在数据驱动决策中日益重要。共形预测提供有限样本下的边际覆盖率,但现有隐私方法常依赖数据切分,降低有效样本量。本文提出一种全数据隐私保护共形预测框架,无需数据切分。该框架利用差分隐私带来的稳定性,控制样本内与样本外共形得分的差距,并结合保守的私有分位数算法以防止覆盖率不足。我们证明通用差分隐私保证可提供统一的覆盖率下界,但通常无法恢复名义上的 $1-α$ 水平。随后通过机制相关的精细化稳定性分析,实现了渐近的名义水平恢复。实验表明,所提方法生成的预测集比基于切分的私有基线更紧致。
原文摘要 · Abstract (English)
Privacy protection and uncertainty quantification are increasingly important in data-driven decision making. Conformal prediction provides finite-sample marginal coverage, but existing private approaches often rely on data splitting, reducing the effective sample size. We propose a full-data privacy-preserving conformal prediction framework that avoids splitting. Our framework leverages stability induced by differential privacy to control the gap between in-sample and out-of-sample conformal scores, and pairs this with a conservative private quantile routine designed to prevent under-coverage. We show that a generic differential privacy guarantee yields a universal coverage floor, yet cannot generally recover the nominal $1-α$ level. We then provide a refined, mechanism-specific stability analysis and yields asymptotic recovery of the nominal level. Experiments demonstrate sharper prediction sets than the split-based private baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。