无需计算海森矩阵,用采样法高效定位关键训练数据
Bayesian Influence Functions for Hessian-Free Data Attribution
- 用随机梯度马尔可夫链蒙特卡洛替代海森矩阵求逆
- 在超大规模神经网络上实现高阶参数交互建模
- 适合需要精准数据溯源的模型调试与可信AI场景
经典影响函数在深度神经网络中面临海森矩阵不可逆和参数空间维数过高的挑战。本文提出局部贝叶斯影响函数(BIF),将海森矩阵求逆替换为可通过随机梯度马尔可夫链蒙特卡洛采样估计的损失曲面统计量。该无海森方法能捕捉参数间的高阶相互作用,并可高效扩展至含数十亿参数的神经网络。实验表明,该方法在预测再训练结果方面达到当前最优性能。
原文摘要 · Abstract (English)
Classical influence functions face significant challenges when applied to deep neural networks, primarily due to non-invertible Hessians and high-dimensional parameter spaces. We propose the local Bayesian influence function (BIF), an extension of classical influence functions that replaces Hessian inversion with loss landscape statistics that can be estimated via stochastic-gradient MCMC sampling. This Hessian-free approach captures higher-order interactions among parameters and scales efficiently to neural networks with billions of parameters. We demonstrate state-of-the-art results on predicting retraining experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。