用损失核分析神经网络如何区分数据,揭示图像语义结构。
The Loss Kernel: A Geometric Probe for Deep Learning Interpretability
- 通过扰动参数计算样本损失的协方差,构建损失相似性度量。
- 在ImageNet上揭示出与WordNet语义层次一致的数据结构。
- 适合关注模型可解释性与数据归属的科研人员。
我们提出损失核,一种用于衡量训练后神经网络中数据点间相似性的可解释性方法。该核是基于低损失保持参数扰动分布下逐样本损失的协方差矩阵。首先在合成多任务问题上验证,结果符合理论预测,能按任务分离输入。随后应用于Inception-v1模型,可视化ImageNet数据结构,发现其与WordNet语义层次高度一致。这证明损失核是一种实用的可解释性工具,可用于数据归因。
原文摘要 · Abstract (English)
We introduce the loss kernel, an interpretability method for measuring similarity between data points according to a trained neural network. The kernel is the covariance matrix of per-sample losses computed under a distribution of low-loss-preserving parameter perturbations. We first validate our method on a synthetic multitask problem, showing it separates inputs by task as predicted by theory. We then apply this kernel to Inception-v1 to visualize the structure of ImageNet, and we show that the kernel's structure aligns with the WordNet semantic hierarchy. This establishes the loss kernel as a practical tool for interpretability and data attribution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。