arXiv:2508.06776cs.LGcs.AI2025-08

通过分析模型激活的零空间,无监督检测大语言模型的表征漂移。

Zero-Direction Probing: A Linear-Algebraic Framework for Deep Analysis of Large-Language-Model Drift

  • 基于线性代数构建理论框架,利用激活的零空间探测模型变化。
  • 提出谱零泄漏度量,在高斯假设下给出漂移的先验阈值。
  • 适用于需要无监督监控模型演化的研究者和系统部署人员。

我们提出零方向探查(ZDP),一个仅依赖理论的框架,用于在无需任务标签或输出评估的情况下,通过Transformer模型激活的零方向检测模型漂移。在假设A1--A6成立的前提下,我们证明了:(i) 方差-泄漏定理,(ii) Fisher零守恒性质,(iii) 低秩更新的秩-泄漏界,以及 (iv) 在线零空间追踪器的对数遗憾保证。我们推导出一种谱零泄漏(SNL)度量,并给出了非渐近尾部界限与浓度不等式,从而在高斯零模型下获得漂移的先验阈值。这些结果表明,监测层激活的左右零空间及其Fisher几何,可提供关于表征变化的明确、可验证的保障。

原文摘要 · Abstract (English)

We present Zero-Direction Probing (ZDP), a theory-only framework for detecting model drift from null directions of transformer activations without task labels or output evaluations. Under assumptions A1--A6, we prove: (i) the Variance--Leak Theorem, (ii) Fisher Null-Conservation, (iii) a Rank--Leak bound for low-rank updates, and (iv) a logarithmic-regret guarantee for online null-space trackers. We derive a Spectral Null-Leakage (SNL) metric with non-asymptotic tail bounds and a concentration inequality, yielding a-priori thresholds for drift under a Gaussian null model. These results show that monitoring right/left null spaces of layer activations and their Fisher geometry provides concrete, testable guarantees on representational change.

模型漂移零空间分析理论框架大模型监控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。