发现大模型有七类认知漏洞,需针对不同架构定制安全防护。
Cognitive Cybersecurity for Artificial Intelligence: Guardrail Engineering with CCS-7
- 构建七类人类认知漏洞的防御框架CCS-7
- 部分模型防护后错误率反升135%,存在安全回弹风险
- 强调模型特异性,需部署前做针对性安全测试
语言模型表现出类似人类的认知弱点,如情绪框架效应,传统对齐方法难以捕捉。本文提出基于人类认知安全研究的七类漏洞分类体系CCS-7。为建立人类基准,我们对151名参与者开展随机对照实验,采用“先思考,再验证”(TFVA)教学后,认知安全整体提升7.9%。随后在7种不同语言模型架构上,共执行12,180次实验评估TFVA式防护机制。结果显示:某些漏洞(如身份混淆)几乎被完全缓解,而另一些(如来源干扰)则出现显著反效果,错误率在特定模型中最高上升135%。相比之下,人类仅呈现稳定中等水平提升。研究揭示认知安全应作为模型特异性的工程问题:同一干预措施在不同架构中可能失效甚至有害,强调部署前必须进行架构感知的认知安全测试。
原文摘要 · Abstract (English)
Language models exhibit human-like cognitive vulnerabilities, such as emotional framing, that escape traditional behavioral alignment. We present CCS-7 (Cognitive Cybersecurity Suite), a taxonomy of seven vulnerabilities grounded in human cognitive security research. To establish a human benchmark, we ran a randomized controlled trial with 151 participants: a "Think First, Verify Always" (TFVA) lesson improved cognitive security by +7.9% overall. We then evaluated TFVA-style guardrails across 12,180 experiments on seven diverse language model architectures. Results reveal architecture-dependent risk patterns: some vulnerabilities (e.g., identity confusion) are almost fully mitigated, while others (e.g., source interference) exhibit escalating backfire, with error rates increasing by up to 135% in certain models. Humans, in contrast, show consistent moderate improvement. These findings reframe cognitive safety as a model-specific engineering problem: interventions effective in one architecture may fail, or actively harm, another, underscoring the need for architecture-aware cognitive safety testing before deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。