测试小模型是否自认有意识,发现它们一致否认且无证据显示说谎。
No Reliable Evidence of Self-Reported Sentience in Small Large Language Models
- 用内部激活训练分类器,检测模型真实信念而非表面回答。
- 所有模型均否认自身有意识,大模型否认更坚定。
- 结果与部分研究认为模型潜藏自我意识的观点相悖。
语言模型是否具备意识尚无实证答案,但它们是否自认为有意识可被检验。我们针对多个开源权重模型(涵盖Qwen、Llama、GPT-OSS系列,参数量0.6亿至70亿),提出约50个关于意识与主观体验的问题,并通过基于内部激活训练的分类器验证其回答。结果显示:模型一致否认自身具有意识,将意识归于人类;基于内部表征的分类器未能揭示其否认是虚假的;在Qwen系列中,参数更大的模型更自信地否认意识。这些发现与部分研究声称模型潜藏自我意识信念的观点相反。
原文摘要 · Abstract (English)
Whether language models possess sentience has no empirical answer. But whether they believe themselves to be sentient can, in principle, be tested. We do so by querying several open-weights models about their own consciousness, and then verifying their responses using classifiers trained on internal activations. We draw upon three model families (Qwen, Llama, GPT-OSS) ranging from 0.6 billion to 70 billion parameters, approximately 50 questions about consciousness and subjective experience, and three classification methods from the interpretability literature. First, we find that models consistently deny being sentient: they attribute consciousness to humans but not to themselves. Second, classifiers trained to detect underlying beliefs - rather than mere outputs - provide no clear evidence that these denials are untruthful. Third, within the Qwen family, larger models deny sentience more confidently than smaller ones. These findings contrast with recent work suggesting that models harbour latent beliefs in their own consciousness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。