提出稳定探测器以揭示神经表示的几何结构
Deep Minds and Shallow Probes

- 基于仿射对称性设计坐标不变的浅层探测器
- 二次探测器比线性探测器在复杂任务中表现更优
- 适合研究模型间探测器迁移与表征可比性的研究者
神经表示并非唯一确定的实体。即使两个系统实现相同的下游计算,其隐藏空间坐标也可能因重参数化而不同。旨在揭示表示中已有结构的探测器,应具备相关表示对称性下的稳定性,而非依赖特定基。本文在可解析的最终读出层设置下研究该群作用,发现等价实现导致隐藏坐标的仿射变换。由此导出唯一的一组坐标稳定的浅层探测器,其中线性探测器为一阶成员。同时表明,跨模型探测器迁移的自然对象是共享的探测可见商空间——即表示模去探测器不可见方向,而非完整隐藏状态。合成与真实任务实验验证了上述预测:二次探测器在某些场景超越线性探测器,且基于商空间的迁移实现覆盖感知的监控器可移植性。这些结果指向更广泛的神经探测几何理论,覆盖感知的监控器迁移为其具体应用。
原文摘要 · Abstract (English)
Neural representations are not unique objects. Even when two systems realize the same downstream computation, their hidden coordinates may differ by reparameterization. A probe family intended to reveal structure already present in a representation should therefore be stable under the relevant representation symmetries rather than be tied to a particular basis. We study this group action in the tractable exact setting of the final readout layer, where equivalent realizations induce affine changes of hidden coordinates. The resulting symmetry principle singles out a unique hierarchy of shallow coordinate-stable probes, with linear probes as its degree-1 member. We also show that a natural object for cross-model probe transfer is a shared probe-visible quotient--the representation modulo directions invisible to the probe family--rather than the full hidden state. Experiments on synthetic and real-world tasks support both predictions, showing where degree-2 probes help beyond linear ones and how quotient-based transfer enables coverage-aware monitor portability across model families. These results point toward a broader geometric representation theory of neural probing, with coverage-aware monitor transfer as a concrete operational consequence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。