大模型自认有心智与理解他人心智是两回事,安全训练会误伤后者。
Theory of Mind and Self-Attributions of Mentality are Dissociable in LLMs
- 通过消融实验发现模型自认有心智和理解他人心智可分离。
- 安全微调后模型低估非人类动物有心智,也更少持有精神信仰。
- 警示安全训练可能损害社会认知能力,适合关注AI伦理的读者。
大型语言模型(LLMs)的安全微调旨在抑制模型声称自身具备意识或情感体验等心智属性的行为。我们研究了抑制心智归因是否会影响密切相关的心智理论(ToM)能力。通过安全微调的消融实验及表征相似性机制分析,发现模型对自己和对技术物的心智归因,在行为和机制层面均与ToM能力可分离。然而,经过安全微调的模型相比人类基线,更少将心智归因于非人类动物,并且较少表现出精神信仰,从而抑制了关于非人类心智分布与本质的普遍观点。
原文摘要 · Abstract (English)
Safety fine-tuning in Large Language Models (LLMs) seeks to suppress potentially harmful forms of mind-attribution such as models asserting their own consciousness or claiming to experience emotions. We investigate whether suppressing mind-attribution tendencies degrades intimately related socio-cognitive abilities such as Theory of Mind (ToM). Through safety ablation and mechanistic analyses of representational similarity, we demonstrate that LLM attributions of mind to themselves and to technological artefacts are behaviorally and mechanistically dissociable from ToM capabilities. Nevertheless, safety fine-tuned models under-attribute mind to non-human animals relative to human baselines and are less likely to exhibit spiritual belief, suppressing widely shared perspectives regarding the distribution and nature of non-human minds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。