arXiv:2510.26717cs.LGcs.DS2025-10

提出一种纯隐私下的协方差估计新方法,精度达理论最优。

On Purely Private Covariance Estimation

  • 通过简单扰动机制实现纯差分隐私下的协方差矩阵发布。
  • 在大数据下误差达弗罗比尼乌斯与谱范数最优,小数据下改进已有界。
  • 适合对隐私与精度要求高的统计分析场景,如医疗数据建模。

我们提出一种针对 $d$ 维协方差矩阵 $Σ$ 的简单扰动机制,在纯差分隐私下实现数据发布。当样本量 $n \geq d^2/\varepsilon$ 时,该机制在弗罗比尼乌斯范数下达到 \\cite{nikolov2023private} 的理论最优误差,同时在所有 $p$-Schatten 范数($p \in [1,\infty]$)中表现最佳;尤其对于 $p \geq 2$,误差为信息论最优,首次实现谱范数下的纯隐私最优。在小样本 $n < d^2/\varepsilon$ 时,通过将输出投影至适当半径的核范数球,可实现最优弗罗比尼乌斯误差 $O(\sqrt{d\;\text{Tr}(Σ)/n})$,优于 \\cite{nikolov2023private} 的 $O(\sqrt{d/n})$ 和 \\cite{dong2022differentially} 的 ${O}(d^{3/4}\sqrt{\text{Tr}(Σ)/n})$。

原文摘要 · Abstract (English)

We present a simple perturbation mechanism for the release of $d$-dimensional covariance matrices $Σ$ under pure differential privacy. For large datasets with at least $n\geq d^2/\varepsilon$ elements, our mechanism recovers the provably optimal Frobenius norm error guarantees of \cite{nikolov2023private}, while simultaneously achieving best known error for all other $p$-Schatten norms, with $p\in [1,\infty]$. Our error is information-theoretically optimal for all $p\ge 2$, in particular, our mechanism is the first purely private covariance estimator that achieves optimal error in spectral norm. For small datasets $n< d^2/\varepsilon$, we further show that by projecting the output onto the nuclear norm ball of appropriate radius, our algorithm achieves the optimal Frobenius norm error $O(\sqrt{d\;\text{Tr}(Σ) /n})$, improving over the known bounds of $O(\sqrt{d/n})$ of \cite{nikolov2023private} and ${O}\big(d^{3/4}\sqrt{\text{Tr}(Σ)/n}\big)$ of \cite{dong2022differentially}.

隐私计算协方差估计差分隐私统计推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。