arXiv:2508.15839cs.CRcs.AI2025-08被引 1

为大模型推理漏洞设计认知安全评估框架,防范看似合法却误导决策的攻击。

CIA+TA Risk Assessment for AI Reasoning Vulnerabilities

  • 构建CIA+TA五维安全框架,新增信任与自主性保障机制
  • 基于实证数据提出量化风险评估方法,揭示防御效果差异超200%
  • 适用于高风险决策场景的大模型系统,需部署认知渗透测试

随着人工智能系统在关键决策中作用增强,其面临利用推理机制而非技术基础设施的新型威胁。本文提出认知网络安全框架,系统性保护AI推理过程免受对抗性操控。贡献有三:首先,确立认知网络安全作为传统网络安全与AI安全的补充领域,应对合法输入扭曲推理却规避常规检测的问题;其次,提出CIA+TA模型,在传统保密性、完整性、可用性基础上增加信任(认知验证)与自主性(人类决策权保留),以适应生成知识主张并参与决策的系统需求;第三,建立基于实证数据的定量风险评估方法,使用可量化的系数帮助组织度量认知安全风险。该框架与OWASP LLM Top 10及MITRE ATLAS对齐,便于落地应用。通过151名参与者和12,180次AI试验的验证发现,相同防御措施在不同架构下效果差异巨大——漏洞减少幅度从96%到放大135%不等,因此必须将认知渗透测试纳入可信AI部署的前置治理要求。

原文摘要 · Abstract (English)

As AI systems increasingly influence critical decisions, they face threats that exploit reasoning mechanisms rather than technical infrastructure. We present a framework for cognitive cybersecurity, a systematic protection of AI reasoning processes from adversarial manipulation. Our contributions are threefold. First, we establish cognitive cybersecurity as a discipline complementing traditional cybersecurity and AI safety, addressing vulnerabilities where legitimate inputs corrupt reasoning while evading conventional controls. Second, we introduce the CIA+TA, extending traditional Confidentiality, Integrity, and Availability triad with Trust (epistemic validation) and Autonomy (human agency preservation), requirements unique to systems generating knowledge claims and mediating decisions. Third, we present a quantitative risk assessment methodology with empirically-derived coefficients, enabling organizations to measure cognitive security risks. We map our framework to OWASP LLM Top 10 and MITRE ATLAS, facilitating operational integration. Validation through previously published studies (151 human participants; 12,180 AI trials) reveals strong architecture dependence: identical defenses produce effects ranging from 96% reduction to 135% amplification of vulnerabilities. This necessitates pre-deployment Cognitive Penetration Testing as a governance requirement for trustworthy AI deployment.

认知安全大模型防护风险评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。