提出智能体认知退化新漏洞及实时防护框架。
QSAF: A Novel Mitigation Framework for Cognitive Degradation in Agentic AI
- 构建六阶段认知退化生命周期模型,识别内部失效机制。
- 部署七项运行时控制,实现内存与逻辑异常的主动干预。
- 借鉴神经科学,适合开发高可靠性智能体系统的团队使用。
我们提出认知退化作为智能体人工智能系统中的一类新型脆弱性。不同于传统的外部对抗威胁(如提示注入),此类故障源于内部问题,包括记忆饥饿、规划器递归、上下文泛滥和输出抑制。这些系统性缺陷导致智能体出现无声漂移、逻辑崩溃和持续幻觉。为此,我们提出Qorvex安全人工智能行为与认知韧性框架(QSAF Domain 10),一个基于六阶段认知退化生命周期的全周期防御框架。该框架包含七项运行时控制(QSAF-BC-001至BC-007),可实时监控智能体子系统,并通过降级路由、饥饿检测和内存完整性保障实现主动缓解。借鉴认知神经科学,我们将智能体架构映射为人类类比,实现对疲劳、饥饿和角色崩塌的早期预警。通过建立正式生命周期与实时缓解机制,本工作确立认知退化为关键新类AI系统漏洞,并提出首个跨平台的智能体鲁棒行为防御模型。
原文摘要 · Abstract (English)
We introduce Cognitive Degradation as a novel vulnerability class in agentic AI systems. Unlike traditional adversarial external threats such as prompt injection, these failures originate internally, arising from memory starvation, planner recursion, context flooding, and output suppression. These systemic weaknesses lead to silent agent drift, logic collapse, and persistent hallucinations over time. To address this class of failures, we introduce the Qorvex Security AI Framework for Behavioral & Cognitive Resilience (QSAF Domain 10), a lifecycle-aware defense framework defined by a six-stage cognitive degradation lifecycle. The framework includes seven runtime controls (QSAF-BC-001 to BC-007) that monitor agent subsystems in real time and trigger proactive mitigation through fallback routing, starvation detection, and memory integrity enforcement. Drawing from cognitive neuroscience, we map agentic architectures to human analogs, enabling early detection of fatigue, starvation, and role collapse. By introducing a formal lifecycle and real-time mitigation controls, this work establishes Cognitive Degradation as a critical new class of AI system vulnerability and proposes the first cross-platform defense model for resilient agentic behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。