用敏感度分析神经网络对数据扰动的响应,揭示模型结构与数据模式的关系。
Susceptibilities and Patterning: A Primer on Linear Response in Bayesian Learning

- 通过后验期望导数定义敏感度,关联数据扰动与模型变化
- 敏感度矩阵可线性求解数据扰动以实现指定结构改变
- 适用于理解模型决策机制,适合研究者和工程师
本文介绍在[arXiv:2504.18274, arXiv:2601.12703]中发展的敏感度理论,用于解释神经网络行为。可观测量ϕ对数据扰动的敏感度定义为后验期望的导数,根据涨落-耗散定理,等于后验协方差。选择不同的ϕ可得到不同对象:逐样本损失对应影响矩阵(即[arXiv:2509.26544]中的贝叶斯影响函数),而局部化分量的可观测量则给出结构敏感度矩阵,将模型组件与数据模式配对。该敏感度矩阵(乘以nβ因子)是数据分布到结构坐标的映射的雅可比矩阵;其伪逆提供了[arXiv:2601.13548]中“模式生成问题”的线性化解法——寻找能引发期望结构变化的数据扰动。理论从统计力学基础出发,详细阐述敏感度、其经验估计器及其与损失曲面几何的关系。
原文摘要 · Abstract (English)
These notes introduce the theory of susceptibilities as developed in [arXiv:2504.18274, arXiv:2601.12703] for interpreting neural networks. The susceptibility of an observable $ϕ$ to a data perturbation is defined as a derivative of a posterior expectation, which by the fluctuation--dissipation theorem equals a posterior covariance. Different choices of $ϕ$ yield different objects: per-sample losses give the influence matrix (the Bayesian influence function of [arXiv:2509.26544]), while component-localized observables give the structural susceptibility matrix that pairs model components with data patterns. The susceptibility matrix is (up to a factor of $nβ$) the Jacobian of the map from data distributions to structural coordinates; its pseudo-inverse provides a linearized solution to the patterning problem of [arXiv:2601.13548]: finding data perturbations that produce a desired structural change. We motivate the theory from its statistical-mechanical foundations, then give a detailed exposition of susceptibilities, their empirical estimators, and their connection to the geometry of the loss landscape.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。