arXiv:2601.18546cs.LG2026-01

仅用梯度就能恢复数据协方差,让优化更高效

Information Hidden in Gradients of Regression with Target Noise

  • 通过注入特定噪声使目标噪声方差等于批量大小,梯度协方差逼近海森矩阵
  • 理论证明在亚高斯输入下,梯度协方差可无偏估计数据协方差Σ
  • 适用于分布式训练、对抗风险估计等场景,实用且鲁棒

二阶信息(如曲率或数据协方差)对优化、诊断和鲁棒性至关重要。然而在许多现代设置中,仅可观测梯度。本文证明,仅凭梯度即可恢复海森矩阵,其等于线性回归中的数据协方差Σ。关键洞察是简单的方差校准:注入高斯噪声,使总目标噪声方差等于批量大小n,此时经验梯度协方差可近似海森矩阵,即使远离最优解。我们在亚高斯输入下提供非渐近的算子范数保证。若无此校准,恢复误差可能达Ω(1)量级。所提方法具实用性(“设目标噪声方差为n”规则)且鲁棒(方差O(n)即能按比例恢复Σ)。应用包括加速优化的预条件、对抗风险估计及仅梯度训练,如分布式系统。实验验证了合成与真实数据上的理论结果。

原文摘要 · Abstract (English)

Second-order information -- such as curvature or data covariance -- is critical for optimisation, diagnostics, and robustness. However, in many modern settings, only the gradients are observable. We show that the gradients alone can reveal the Hessian, equalling the data covariance $Σ$ for the linear regression. Our key insight is a simple variance calibration: injecting Gaussian noise so that the total target noise variance equals the batch size ensures that the empirical gradient covariance closely approximates the Hessian, even when evaluated far from the optimum. We provide non-asymptotic operator-norm guarantees under sub-Gaussian inputs. We also show that without such calibration, recovery can fail by an $Ω(1)$ factor. The proposed method is practical (a "set target-noise variance to $n$" rule) and robust (variance $\mathcal{O}(n)$ suffices to recover $Σ$ up to scale). Applications include preconditioning for faster optimisation, adversarial risk estimation, and gradient-only training, for example, in distributed systems. We support our theoretical results with experiments on synthetic and real data.

优化梯度分析协方差恢复线性回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。