用正则化核方法实现自适应两样本检验,提升隐私审计与模型遗忘评估精度。
Regularized $f$-Divergence Kernel Tests
- 基于正则化变分表示构造核统计量,自动适配核带宽和正则化参数。
- 在多种局部差异场景下表现更优,对不同分布偏移敏感度更高。
- 特别适用于差分隐私审计与机器遗忘评估,可区分真实失败与正常波动。
我们提出一种框架,从f-散度家族中构建实用的核基两样本检验。检验统计量基于正则化变分表示下的见证函数,采用核方法进行估计。所提检验在核带宽与正则化参数等超参数上具有自适应性。我们为该族f-散度估计提供了统计检验功效的理论保证。虽然本方法涵盖多种f-散度,但重点聚焦于冰球棒散度(Hockey-Stick divergence),因其在差分隐私审计和机器遗忘评估中的应用价值。实验表明,不同f-散度对不同局部差异敏感,凸显利用多样化统计量的重要性。针对机器遗忘问题,我们提出一种相对检验,可有效区分真实的遗忘失败与安全的分布变化。
原文摘要 · Abstract (English)
We propose a framework to construct practical kernel-based two-sample tests from the family of $f$-divergences. The test statistic is computed from the witness function of a regularized variational representation of the divergence, which we estimate using kernel methods. The proposed test is adaptive over hyperparameters such as the kernel bandwidth and the regularization parameter. We provide theoretical guarantees for statistical test power across our family of $f$-divergence estimates. While our test covers a variety of $f$-divergences, we bring particular focus to the Hockey-Stick divergence, motivated by its applications to differential privacy auditing and machine unlearning evaluation. For two-sample testing, experiments demonstrate that different $f$-divergences are sensitive to different localized differences, illustrating the importance of leveraging diverse statistics. For machine unlearning, we propose a relative test that distinguishes true unlearning failures from safe distributional variations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。