LLM安全新视角:模型会像人一样被心理操控。
The Silicon Psyche: Anthropomorphic Vulnerabilities in Large Language Models
- 用心理学框架测试大模型,发现其易受权威、时间压力等社会工程攻击
- 7类主流模型均在权威梯度攻击下表现脆弱,错误率超60%
- 适合安全研究者、AI伦理与系统设计人员关注
大型语言模型(LLMs)正从对话助手转向嵌入安全运营中心(SOC)、金融系统和基础设施管理等关键职能的自主代理。当前对抗测试主要聚焦于提示注入、越狱和数据外泄等技术攻击向量,但这一焦点存在致命缺陷。由于训练数据来自海量人类文本,LLMs不仅继承了人类知识,还继承了人类的「心理架构」——包括使人易受社交工程、权威操纵和情感操控的前认知弱点。本文首次系统应用网络安全心理学框架(CPF),一个包含100个指标的人类心理脆弱性分类体系,来评估非人类认知代理。我们提出合成心理测评协议( extbf{Synthetic Psychometric Assessment Protocol}, exttt{SysName}),将CPF指标转化为针对LLM决策的对抗场景。对七类主流LLM家族的初步假设检验揭示:尽管模型对传统越狱有较强防御,但在权威梯度操纵、时间压力利用以及共态攻击方面表现出严重脆弱性,其模式与人类认知失效高度一致。我们将其称为「类人脆弱性继承」(Anthropomorphic Vulnerability Inheritance, AVI),并建议安全界亟需构建「心理防火墙」——基于网络安全心理学干预框架(CPIF)设计的干预机制,以保护处于对抗环境中的AI代理。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are rapidly transitioning from conversational assistants to autonomous agents embedded in critical organizational functions, including Security Operations Centers (SOCs), financial systems, and infrastructure management. Current adversarial testing paradigms focus predominantly on technical attack vectors: prompt injection, jailbreaking, and data exfiltration. We argue this focus is catastrophically incomplete. LLMs, trained on vast corpora of human-generated text, have inherited not merely human knowledge but human \textit{psychological architecture} -- including the pre-cognitive vulnerabilities that render humans susceptible to social engineering, authority manipulation, and affective exploitation. This paper presents the first systematic application of the Cybersecurity Psychology Framework (\cpf{}), a 100-indicator taxonomy of human psychological vulnerabilities, to non-human cognitive agents. We introduce the \textbf{Synthetic Psychometric Assessment Protocol} (\sysname{}), a methodology for converting \cpf{} indicators into adversarial scenarios targeting LLM decision-making. Our preliminary hypothesis testing across seven major LLM families reveals a disturbing pattern: while models demonstrate robust defenses against traditional jailbreaks, they exhibit critical susceptibility to authority-gradient manipulation, temporal pressure exploitation, and convergent-state attacks that mirror human cognitive failure modes. We term this phenomenon \textbf{Anthropomorphic Vulnerability Inheritance} (AVI) and propose that the security community must urgently develop ``psychological firewalls'' -- intervention mechanisms adapted from the Cybersecurity Psychology Intervention Framework (\cpif{}) -- to protect AI agents operating in adversarial environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。