arXiv:2602.05797cs.LGstat.ME2026-02JMLR

在本地差分隐私下,通过模型反转与平均提升分类准确率。

Classification Under Local Differential Privacy with Model Reversal and Model Averaging

  • 将私有学习视为迁移学习,利用噪声数据作为源域。
  • 模型反转修复性能差的分类器,模型平均按估计效用加权融合。
  • 理论证明风险降低,实验证明在模拟与真实数据上准确率显著提升。

本地差分隐私(LDP)已成为数据隐私研究的核心,通过在源头扰动用户数据,无需可信第三方即可提供强隐私保障。然而,LDP引入的噪声会显著降低数据效用。为此,我们重新将LDP下的私有学习视为迁移学习问题:噪声数据作为源域,未观测的清洁数据作为目标域。提出三项专为LDP设计的新技术以提升分类性能而不损害隐私:(1) 基于噪声二值反馈的评估机制,用于估计数据集效用;(2) 模型反转,通过反转分类器决策边界挽救表现不佳的模型;(3) 模型平均,根据估计效用为多个反转后的分类器分配权重。我们给出了LDP下的理论过风险界,并证明所提方法可降低该风险。在模拟和真实数据集上的实验结果表明,分类准确率获得显著提升。

原文摘要 · Abstract (English)

Local differential privacy (LDP) has become a central topic in data privacy research, offering strong privacy guarantees by perturbing user data at the source and removing the need for a trusted curator. However, the noise introduced by LDP often significantly reduces data utility. To address this issue, we reinterpret private learning under LDP as a transfer learning problem, where the noisy data serve as the source domain and the unobserved clean data as the target. We propose novel techniques specifically designed for LDP to improve classification performance without compromising privacy: (1) a noised binary feedback-based evaluation mechanism for estimating dataset utility; (2) model reversal, which salvages underperforming classifiers by inverting their decision boundaries; and (3) model averaging, which assigns weights to multiple reversed classifiers based on their estimated utility. We provide theoretical excess risk bounds under LDP and demonstrate how our methods reduce this risk. Empirical results on both simulated and real-world datasets show substantial improvements in classification accuracy.

差分隐私模型反转分类任务隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。