arXiv:2605.28823cs.CL2026-05中稿 · the 6th Workshop o…被引 1

开发可检测大模型思考内容的低成本探针,实现对抽象概念的追踪与监控。

What are They Thinking? Delineation, Probing, and Tracking of Concepts in LLMs

论文配图:What are They Thinking? Delineation, Probing, and Tracking of Concepts in LLMs
图 1 · 摘自论文原文
  • 通过构建正反例数据集,精准定义抽象概念边界。
  • 在多层大模型中训练线性探针,成功识别四类概念。
  • 探针可跨长上下文追踪概念变化,适用于多种模型。

随着大语言模型影响扩大,理解其决策机制变得至关重要。一种有效方式是开发探针,检测模型嵌入向量中是否存在广泛且高层次的抽象概念——这正是我们所说的模型‘思考’的内容。这些探针需成本低、易部署,以便在正常运行中监控多种概念。本文首次系统推进该能力:首先,通过构建含概念存在与不存在的语料库,精确界定抽象概念;其次,在三个不同大模型的多层中训练并测试一系列线性探针,探索所需探针复杂度;最后,验证探针可在更长上下文中持续追踪概念。实验涵盖四个概念和三种模型。未来若扩展至更多概念,将实现对新模型的全面监测。

原文摘要 · Abstract (English)

As the influence of LLMs expands, it is imperative to gain insight into their decisions. One way to do that is to develop probes that detect the presence or absence of a broad set of high-level abstract concepts within the embeddings computed in an LLM - which is what we might say a model is ``thinking" about. Such probes should be low-cost and easily applicable to any LLM, so that monitoring for many concepts is possible during normal operation. In this paper, we take the first steps towards developing the capability of creating many such probes by defining and executing examples of the key tasks needed: first, the careful delineation of a high-level abstract concept through the creation of a dataset with the concept both present and then absent. Then, the training and testing of a set of linear probes to detect the concept on any layer of an LLM, including an exploration of the complexity of the probe needed. Finally, we show that such probes can track concepts across larger contexts. This is done with four separate concepts and three different LLMs. When this process is scaled to many more concepts, it will create the ability to monitor new models.

大模型解释概念探针抽象理解可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。