揭示线性探针失效的谱机制,提供可验证的稳定性判断准则。
Spectral Identifiability for Interpretable Probe Geometry
- 基于谱间隙与费舍尔误差的对比,提出可验证的探针稳定性判据。
- 当谱间隙大于费舍尔估计误差时,探针准确率稳定;否则呈相变式崩溃。
- 适用于需要可靠评估神经表征的科研人员,尤其关注探针可信度者。
线性探针广泛用于解释和评估神经表征,但其可靠性存疑:某些情况下表现良好,却在其他情形下突然失效。本文揭示了这一现象背后的谱机制,提出谱可辨识性原则(SIP),一种类费舍尔的可验证条件,用于判断探针稳定性。当分离任务相关方向的谱间隙大于费舍尔估计误差时,估计子空间集中,准确率保持一致;而谱间隙缩小则引发相变式的不稳定性。通过有限样本分析,将谱间隙几何、样本量与误分类风险关联起来,提供可解释的诊断工具,而非松散的泛化界。控制性合成实验中,精确计算费舍尔量,验证了预测效果,表明谱检查可在探针扭曲下游评估前预判其不可靠性。
原文摘要 · Abstract (English)
Linear probes are widely used to interpret and evaluate neural representations, yet their reliability remains unclear, as probes may appear accurate in some regimes but collapse unpredictably in others. We uncover a spectral mechanism behind this phenomenon and formalize it as the Spectral Identifiability Principle (SIP), a verifiable Fisher-inspired condition for probe stability. When the eigengap separating task-relevant directions is larger than the Fisher estimation error, the estimated subspace concentrates and accuracy remains consistent, whereas closing this gap induces instability in a phase-transition manner. Our analysis connects eigengap geometry, sample size, and misclassification risk through finite-sample reasoning, providing an interpretable diagnostic rather than a loose generalization bound. Controlled synthetic studies, where Fisher quantities are computed exactly, confirm these predictions and show how spectral inspection can anticipate unreliable probes before they distort downstream evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。