arXiv:2604.17698cs.LGcs.CL2026-04被引 3

用几何稳定性同时预测模型可控性与检测性能退化

The Geometric Canary: Predicting Steerability and Detecting Drift via Representational Stability

  • 通过任务对齐的几何稳定性衡量模型可操控性
  • 监督方法预测可控性准确率高达0.89-0.97
  • 无监督稳定性适合部署后监测,误报率低6倍

语言模型可靠部署需具备两种看似不同但共享几何基础的能力:预测模型是否可被定向控制,以及检测其内部结构退化。研究发现,表示的几何稳定性——即表示间距离结构的一致性——能同时解决这两个问题。基于任务对齐的监督式Shesha方法在35-69个嵌入模型和三个NLP任务上,以ρ=0.89-0.97的高相关性精准预测线性可操控性,且捕捉到类别可分性之外的独特方差(偏相关ρ=0.62-0.76)。关键发现:无监督稳定性在真实任务上的操控预测失效(ρ≈0.10),表明任务对齐对可控性预测至关重要;但其在退化检测中表现卓越,在后训练对齐阶段测得的几何变化量比CKA高出近2倍(Llama模型达5.23倍),且在73%模型中提供更早预警,误报率仅为Procrustes的1/6。因此,监督与无监督稳定性共同构成大模型全生命周期诊断工具:前者用于部署前可控性评估,后者用于部署后持续监控。

原文摘要 · Abstract (English)

Reliable deployment of language models requires two capabilities that appear distinct but share a common geometric foundation: predicting whether a model will accept targeted behavioral control, and detecting when its internal structure degrades. We show that geometric stability, the consistency of a representation's pairwise distance structure, addresses both. Supervised Shesha variants that measure task-aligned geometric stability predict linear steerability with near-perfect accuracy ($ρ= 0.89$-$0.97$) across 35-69 embedding models and three NLP tasks, capturing unique variance beyond class separability (partial $ρ= 0.62$-$0.76$). A critical dissociation emerges: unsupervised stability fails entirely for steering on real-world tasks ($ρ\approx 0.10$), revealing that task alignment is essential for controllability prediction. However, unsupervised stability excels at drift detection, measuring nearly $2\times$ greater geometric change than CKA during post-training alignment (up to $5.23\times$ in Llama) while providing earlier warning in 73\% of models and maintaining a $6\times$ lower false alarm rate than Procrustes. Together, supervised and unsupervised stability form complementary diagnostics for the LLM deployment lifecycle: one for pre-deployment controllability assessment, the other for post-deployment monitoring.

模型可操控性几何稳定性退化检测大模型监控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。