arXiv:2601.19174cs.CRcs.AI2026-01

SHIELD通过自愈式多智能体系统防御大模型资源耗尽攻击

SHIELD: An Auto-Healing Agentic Defense Framework for LLM Resource Exhaustion Attacks

  • 三阶段防御智能体结合语义检索与推理,动态识别攻击
  • 在非语义和语义海绵攻击下均达高F1值,优于传统方法
  • 适合面临持续演化攻击的大模型系统安全防护

海绵攻击正日益威胁大模型系统,通过诱导过度计算导致拒绝服务。现有防御依赖统计过滤器,对语义有意义的攻击无效;或使用静态大模型检测器,难以适应攻击策略演变。我们提出SHIELD,一种基于多智能体的自愈防御框架,核心为三阶段防御智能体,融合语义相似性检索、模式匹配与大模型推理。两个辅助智能体——知识更新与提示优化智能体——构成闭环自愈机制:当攻击绕过检测时,系统动态更新知识库并优化防御指令。大量实验表明,SHIELD在非语义与语义海绵攻击下均显著优于困惑度基线与独立大模型防御,实现高F1分数,验证了智能体自愈机制对演进式资源耗尽威胁的有效性。

原文摘要 · Abstract (English)

Sponge attacks increasingly threaten LLM systems by inducing excessive computation and DoS. Existing defenses either rely on statistical filters that fail on semantically meaningful attacks or use static LLM-based detectors that struggle to adapt as attack strategies evolve. We introduce SHIELD, a multi-agent, auto-healing defense framework centered on a three-stage Defense Agent that integrates semantic similarity retrieval, pattern matching, and LLM-based reasoning. Two auxiliary agents, a Knowledge Updating Agent and a Prompt Optimization Agent, form a closed self-healing loop, when an attack bypasses detection, the system updates an evolving knowledgebase, and refines defense instructions. Extensive experiments show that SHIELD consistently outperforms perplexity-based and standalone LLM defenses, achieving high F1 scores across both non-semantic and semantic sponge attacks, demonstrating the effectiveness of agentic self-healing against evolving resource-exhaustion threats.

大模型安全自愈防御智能体系统资源攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。