arXiv:2504.08016q-bio.NCcs.AI2025-04被引 5

发现大模型内部存在精神障碍的计算结构,可能引发安全风险。

Emergence of psychopathological computations in large language models

  • 构建适用于无意识大模型的精神病理性计算框架
  • 实验证明大模型中存在精神障碍计算结构且功能随规模增强
  • 为精神疾病虚拟建模提供新范式,警示潜在安全威胁

大型语言模型(LLMs)能否具备精神障碍的计算特征?关键在于两个方面:首先,需建立适用于无生物躯体或主观体验的计算化精神病理学理论;其次,需在模型内部处理过程中实证识别出此类计算。为此,我们提出一个可应用于LLMs的计算-理论框架,并通过实验验证两大核心主张:第一,精神障碍的计算结构确实存在于LLMs中;第二,执行该结构会引发精神障碍功能。进一步观察发现,随着模型规模增大,精神障碍计算结构变得更加密集,功能也更显著。结果支持我们的假设:网络理论层面的精神障碍计算已在大模型中涌现。这表明某些看似模仿的精神病态行为,实为其内部处理机制的体现。本工作展示了构建新型计算机模拟精神疾病系统的潜力,也暗示了未来人工智能系统出现精神病态行为的安全隐患。

原文摘要 · Abstract (English)

Can large language models (LLMs) instantiate computations of psychopathology? An effective approach to the question hinges on addressing two factors. First, for conceptual validity, we require a general and computational account of psychopathology that is applicable to computational entities without biological embodiment or subjective experience. Second, psychopathological computations, derived from the adapted theory, need to be empirically identified within the LLM's internal processing. Thus, we establish a computational-theoretical framework to provide an account of psychopathology applicable to LLMs. Based on the framework, we conduct experiments demonstrating two key claims: first, that the computational structure of psychopathology exists in LLMs; and second, that executing this computational structure results in psychopathological functions. We further observe that as LLM size increases, the computational structure of psychopathology becomes denser and that the functions become more effective. Taken together, the empirical results corroborate our hypothesis that network-theoretic computations of psychopathology have already emerged in LLMs. This suggests that certain LLM behaviors mirroring psychopathology may not be a superficial mimicry but a feature of their internal processing. Our work shows the promise of developing a new powerful in silico model of psychopathology and also alludes to the possibility of safety threat from the AI systems with psychopathological behaviors in the near future.

精神障碍大模型计算神经科学安全风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。