用马氏余弦相似度更准确预测线性探测器的泛化性能。
Comparing Linear Probes with Mahalanobis Cosine Similarity

- 提出马氏余弦相似度,基于测试数据协方差重加权内积。
- 理论证明在高斯分布下,马氏余弦与OOD AUROC呈线性关系(R²=0.98)。
- 适用于跨模型、跨层、跨概念的探测器比较,尤其适合可解释性研究。
线性探测器广泛用于可解释性研究,通常通过余弦相似度进行比较。马氏余弦相似度(MCS)通过测试数据协方差对内积进行重加权,是一种任务感知的自然优化。Ying 等人(2026)报告称,探测器与在分布外(OOD)数据上训练的参考探测器之间的MCS,几乎完美地线性预测了该探测器的OOD AUROC(R² = 0.98)。本文将这一经验发现扩展至多种模型、层和概念领域,并以闭式形式证明了这一普遍现象:当类别平衡且投影服从高斯分布时,OOD AUROC与探测器到参考探测器的MCS均为信号噪声比(SNR)的sigmoid函数,因此两者呈线性关系。该理论还预测了线性失效的条件,并通过实证验证。马氏余弦相似度为比较线性探测器提供了一种理论合理且实证有效的替代方案,优于欧氏余弦相似度。
原文摘要 · Abstract (English)
Linear probes are widely used in interpretability research and often compared by cosine similarity. The Mahalanobis cosine similarity (MCS) between two directions, which reweights the inner product by test data covariance, is a natural task-aware refinement. Ying et al. (2026) report that a probe's MCS to a reference probe trained on the out-of-distribution (OOD) data near-perfectly linearly predicts the probe's OOD AUROC (R^2 = 0.98). Here, we extend this empirical finding across models, layers, and concept domains, and prove this general phenomenon in closed form: For balanced classes whose projections are Gaussian, OOD AUROC and MCS to the reference probe are linear because both are sigmoid-shaped functions of the probe's signal-to-noise ratio (SNR) on the test data. The theory also predicts when this linearity fails, which we verify empirically. MCS offers a theoretically grounded and empirically effective alternative to Euclidean cosine similarity for comparing linear probes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。