大模型无法真正反思自身语言知识,其自评结果不可靠。
Language Models Fail to Introspect About Their Knowledge of Language
- 用字符串概率衡量模型真实语言能力,对比其自我判断
- 21个开源模型均未表现出超越其他模型的自省能力
- 适合关注大模型可解释性与评估方法的读者
近期研究关注大语言模型(LLMs)是否能反思自身内部状态。此类能力可提升模型可解释性,并支持在语言学中使用标准自省方法(如询问“句子是否语法正确”)来评估模型的语法知识。我们系统研究了21个开源LLM在语法知识和词预测两个领域的自省能力。关键在于,模型的内部语言知识理论上可通过字符串概率直接测量。随后评估模型对元语言提示的回应是否真实反映其内部知识。我们提出一种新自省度量:模型回应能否预测自身字符串概率,且优于具有近似内部知识的另一模型。尽管元语言提示和概率比较均取得高任务准确率,但未发现模型具备特权的“自我访问”能力。通过通用任务、控制模型相似性及覆盖广泛开源模型,我们表明大模型无法自省,并为避免将提示响应误认为语言泛化提供了新证据。
原文摘要 · Abstract (English)
There has been recent interest in whether large language models (LLMs) can introspect about their own internal states. Such abilities would make LLMs more interpretable, and also validate the use of standard introspective methods in linguistics to evaluate grammatical knowledge in models (e.g., asking "Is this sentence grammatical?"). We systematically investigate emergent introspection across 21 open-source LLMs, in two domains where introspection is of theoretical interest: grammatical knowledge and word prediction. Crucially, in both domains, a model's internal linguistic knowledge can be theoretically grounded in direct measurements of string probability. We then evaluate whether models' responses to metalinguistic prompts faithfully reflect their internal knowledge. We propose a new measure of introspection: the degree to which a model's prompted responses predict its own string probabilities, beyond what would be predicted by another model with nearly identical internal knowledge. While both metalinguistic prompting and probability comparisons lead to high task accuracy, we do not find evidence that LLMs have privileged "self-access". By using general tasks, controlling for model similarity, and evaluating a wide range of open-source models, we show that LLMs cannot introspect, and add new evidence to the argument that prompted responses should not be conflated with models' linguistic generalizations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。