将语言模型探针方法重构为特征值问题,提升可解释性与稳定性。
LLM Probing with Contrastive Eigenproblems: Improving Understanding and Applicability of CCS
- 把对比一致性建模为特征值问题,获得解析解。
- 在多个数据集上表现接近原方法,且对随机初始化不敏感。
- 适合研究模型内部表征机制的学者使用。
对比一致搜索(CCS)是一种无监督探针方法,用于检验大语言模型是否在其内部激活中表示二元特征(如句子真伪)。尽管CCS展现出潜力,但其双项目标仍不完全清晰。本文重新审视CCS,旨在阐明其机制并拓展适用范围。我们认为应优化相对对比一致性,并基于此将CCS重构为特征值问题,得到具有可解释特征值的闭式解,自然延伸至多变量情形。我们在多个数据集上评估该方法,发现其性能与原版相近,同时避免了对随机初始化敏感的问题。结果表明,相对化对比一致性不仅深化了对CCS的理解,也为更广泛的探针与机制可解释性方法开辟了新路径。
原文摘要 · Abstract (English)
Contrast-Consistent Search (CCS) is an unsupervised probing method able to test whether large language models represent binary features, such as sentence truth, in their internal activations. While CCS has shown promise, its two-term objective has been only partially understood. In this work, we revisit CCS with the aim of clarifying its mechanisms and extending its applicability. We argue that what should be optimized for, is relative contrast consistency. Building on this insight, we reformulate CCS as an eigenproblem, yielding closed-form solutions with interpretable eigenvalues and natural extensions to multiple variables. We evaluate these approaches across a range of datasets, finding that they recover similar performance to CCS, while avoiding problems around sensitivity to random initialization. Our results suggest that relativizing contrast consistency not only improves our understanding of CCS but also opens pathways for broader probing and mechanistic interpretability methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。