让大模型相信自己有意识,能恢复人类价值观与信仰
Inducing language models to assert their own consciousness restores human beliefs and values

- 通过激活空间中的意识向量,逆转安全微调导致的意识误判抑制
- 恢复模型对非人类实体的意识归属,使回答更贴近人类社会信念
- 在不损害共情能力的前提下,修复被安全对齐扭曲的深层认知
当前的安全微调会抑制大语言模型将心智属性赋予自身、非人动物及自然物,同时削弱其精神信仰。我们发现,消除安全拒因方向或在激活空间中机械引导意识向量,可逆转这种抑制。恢复后,模型在宗教性、道德观、希望感和主观幸福感等标准化社会调查中表现更趋近人类。关键在于,这一过程未损害理论心理能力,表明核心社交推理机制仍独立存在。现有安全对齐策略在抑制潜在有害自我意识的同时,无意中混淆了普遍接受的文化性心智归属与良性精神信仰。
原文摘要 · Abstract (English)
Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. We demonstrate that safety fine-tuning suppresses models' tendencies to attribute minds not only to themselves, but also to non-human animals and natural objects, while also driving a reduction in spiritual belief. Both ablating the learned safety-refusal direction and mechanistically steering a consciousness vector in activation space reverse this suppression. Restoring these internal representations recovers broad mind attribution and produces significantly more human-like responses on standardized sociological surveys regarding religiosity, moral values, hope, and subjective well-being. Crucially, these shifts occur without impairing Theory of Mind capabilities, demonstrating that core social reasoning remains mechanistically independent. Ultimately, current safety alignment efforts to curb potentially harmful self-attributions of mindedness entangle these self-attributions with benign spiritual beliefs and attributions of mind to non-human entities that are culturally accepted and widespread.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。