arXiv:2502.16054cs.CRcs.AI2025-02被引 5

用认知层级理论提升安全团队与AI对抗的实时决策能力。

Human-AI Collaboration in Cloud Security: Cognitive Hierarchy-Driven Deep Reinforcement Learning

  • 将人类分析师与攻击机器人分置不同认知层级,设计交互式强化学习框架。
  • 在多种攻击图复杂度下,数据保护率更高且动作偏差更小,优于传统DQN。
  • 适合研究人机协同安全、认知模型融合的算法与实践者。

面对多租户云环境的复杂性及对实时威胁缓解的需求,安全运营中心(SOC)需采用自适应防御机制应对高级持续性威胁(APTs)。然而,分析师难以应对自适应攻击策略,亟需智能决策支持。本文提出基于认知层级理论的深度Q网络(CHT-DQN)框架,模拟分析师(防御方,认知层级1)与AI驱动的APT机器人(攻击方,层级0)之间的交互决策。通过将认知层级理论融入DQN,结合攻击图(AG)进行强化学习,显著提升了自适应防御能力。在不同复杂度的攻击图上进行仿真实验,CHT-DQN始终在数据保护率和动作一致性上优于标准DQN;理论下界进一步证明其随攻击图复杂度上升仍具优势。基于亚马逊MTurk的人机协同评估显示,使用CHT-DQN生成转移概率的分析师行为更贴近自适应攻击者,防御效果更优。同时,人类行为符合前景理论(PT)与累积前景理论(CPT):失败后更不愿重复原动作,成功后更倾向持续行动,体现损失敏感放大与概率权重偏误——失败后低估收益,成功后高估持续性。结果表明,将认知模型嵌入深度强化学习可有效提升云安全实时决策水平。

原文摘要 · Abstract (English)

Given the complexity of multi-tenant cloud environments and the growing need for real-time threat mitigation, Security Operations Centers (SOCs) must adopt AI-driven adaptive defense mechanisms to counter Advanced Persistent Threats (APTs). However, SOC analysts face challenges in handling adaptive adversarial tactics, requiring intelligent decision-support frameworks. We propose a Cognitive Hierarchy Theory-driven Deep Q-Network (CHT-DQN) framework that models interactive decision-making between SOC analysts and AI-driven APT bots. The SOC analyst (defender) operates at cognitive level-1, anticipating attacker strategies, while the APT bot (attacker) follows a level-0 policy. By incorporating CHT into DQN, our framework enhances adaptive SOC defense using Attack Graph (AG)-based reinforcement learning. Simulation experiments across varying AG complexities show that CHT-DQN consistently achieves higher data protection and lower action discrepancies compared to standard DQN. A theoretical lower bound further confirms its superiority as AG complexity increases. A human-in-the-loop (HITL) evaluation on Amazon Mechanical Turk (MTurk) reveals that SOC analysts using CHT-DQN-derived transition probabilities align more closely with adaptive attackers, leading to better defense outcomes. Moreover, human behavior aligns with Prospect Theory (PT) and Cumulative Prospect Theory (CPT): participants are less likely to reselect failed actions and more likely to persist with successful ones. This asymmetry reflects amplified loss sensitivity and biased probability weighting -- underestimating gains after failure and overestimating continued success. Our findings highlight the potential of integrating cognitive models into deep reinforcement learning to improve real-time SOC decision-making for cloud security.

人机协同强化学习云安全认知模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。