通过扰动中间激活值,揭示模型内部结构与偏差。
APEX: Probing Neural Networks via Activation Perturbation
- 在推理阶段扰动隐藏层输出,保持输入和参数不变。
- 小噪声下可衡量样本规律性,大噪声下暴露模型训练偏见。
- 适合研究模型内在机制与对抗攻击防御的开发者。
现有神经网络探测方法主要依赖输入空间分析或参数扰动,均难以获取中间表示中的结构信息。我们提出激活扰动探索(APEX),一种推理时的探测范式:固定输入与模型参数,仅扰动隐藏激活值。理论上,激活扰动能促使行为从样本相关转向模型相关,抑制输入特异性信号并放大表示层面的结构特征,且输入扰动是该框架下的受限特例。通过典型案例验证,小噪声下APEX提供轻量高效的样本规律性度量,与已有指标一致,并能区分结构化与随机标签模型,揭示语义连贯的预测变化;大噪声下则暴露训练引入的模型级偏差,如后门模型中预测高度集中于目标类别。结果表明,APEX为理解神经网络提供了超越输入空间的新视角。
原文摘要 · Abstract (English)
Prior work on probing neural networks primarily relies on input-space analysis or parameter perturbation, both of which face fundamental limitations in accessing structural information encoded in intermediate representations. We introduce Activation Perturbation for EXploration (APEX), an inference-time probing paradigm that perturbs hidden activations while keeping both inputs and model parameters fixed. We theoretically show that activation perturbation induces a principled transition from sample-dependent to model-dependent behavior by suppressing input-specific signals and amplifying representation-level structure, and further establish that input perturbation corresponds to a constrained special case of this framework. Through representative case studies, we demonstrate the practical advantages of APEX. In the small-noise regime, APEX provides a lightweight and efficient measure of sample regularity that aligns with established metrics, while also distinguishing structured from randomly labeled models and revealing semantically coherent prediction transitions. In the large-noise regime, APEX exposes training-induced model-level biases, including a pronounced concentration of predictions on the target class in backdoored models. Overall, our results show that APEX offers an effective perspective for exploring, and understanding neural networks beyond what is accessible from input space alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。