arXiv:2512.09103cs.LGmath.OC2025-12

提出新型几何鲁棒性数据归因方法,让深度网络的归因结果更可信。

Natural Geometry of Robust Data Attribution: From Convex Models to Deep Networks

  • 用模型自身特征协方差定义自然水桶距离,对抗深层网络的谱放大问题。
  • 在CIFAR-10上对ResNet-18实现68.7%的归因对认证率,远超传统方法的0%。
  • 可检测标签噪声,仅看前20%训练数据就识别94.1%的错误标签。

数据归因方法用于识别影响模型预测的关键训练样本,但其对分布扰动敏感,实用性受限。本文提出从凸模型到深度网络的统一认证鲁棒归因框架。在凸情形下,推导出具有可证明覆盖保证的Wasserstein鲁棒影响函数(W-RIF)。对于深度网络,揭示欧氏认证因谱放大效应而失效——深层表示的固有病态使利普希茨常数膨胀超过10,000倍。这解释了为何标准TRAK得分虽准确却几何脆弱:朴素欧氏鲁棒性分析认证率为0%。核心贡献是引入自然水桶度量,衡量由模型自身特征协方差诱导的几何扰动。该度量消除谱放大,将最坏情况敏感度降低76倍,稳定归因估计。在CIFAR-10与ResNet-18上,自然水桶TRAK认证68.7%的排名对,而欧氏基线为0%,据我们所知为首个非空的神经网络归因认证边界。进一步证明,分析中出现的自影响项等于控制归因稳定性的利普希茨常数,为基于杠杆的异常检测提供理论依据。实验表明,自影响在标签噪声检测中达到0.970 AUROC,仅分析前20%训练数据即可识别94.1%的错误标签。

原文摘要 · Abstract (English)

Data attribution methods identify which training examples are responsible for a model's predictions, but their sensitivity to distributional perturbations undermines practical reliability. We present a unified framework for certified robust attribution that extends from convex models to deep networks. For convex settings, we derive Wasserstein-Robust Influence Functions (W-RIF) with provable coverage guarantees. For deep networks, we demonstrate that Euclidean certification is rendered vacuous by spectral amplification -- a mechanism where the inherent ill-conditioning of deep representations inflates Lipschitz bounds by over $10{,}000\times$. This explains why standard TRAK scores, while accurate point estimates, are geometrically fragile: naive Euclidean robustness analysis yields 0\% certification. Our key contribution is the Natural Wasserstein metric, which measures perturbations in the geometry induced by the model's own feature covariance. This eliminates spectral amplification, reducing worst-case sensitivity by $76\times$ and stabilizing attribution estimates. On CIFAR-10 with ResNet-18, Natural W-TRAK certifies 68.7\% of ranking pairs compared to 0\% for Euclidean baselines -- to our knowledge, the first non-vacuous certified bounds for neural network attribution. Furthermore, we prove that the Self-Influence term arising from our analysis equals the Lipschitz constant governing attribution stability, providing theoretical grounding for leverage-based anomaly detection. Empirically, Self-Influence achieves 0.970 AUROC for label noise detection, identifying 94.1\% of corrupted labels by examining just the top 20\% of training data.

数据归因鲁棒性深度学习异常检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。