arXiv:2603.29062cs.CRcs.AI2026-03被引 4

为政府类AI聊天机器人构建多层防御框架,有效抵御复杂攻击

CivicShield: A Cross-Domain Defense-in-Depth Framework for Securing Government-Facing AI Chatbots Against Multi-Turn Adversarial Attacks

  • 采用七层纵深防御体系,融合零信任与行为检测等多重机制
  • 在1436个攻击场景中实现72.9%的检测率,误报率仅2.9%
  • 适合关注政务AI安全与合规部署的开发者与机构

基于大语言模型的政府服务聊天机器人存在严重安全漏洞。当前防御手段面对多轮对抗攻击成功率超过90%,单层防护也以相似比例被绕过。本文提出CivicShield,一种面向政府类AI聊天机器人的跨域纵深防御框架。该框架融合网络安全、形式化验证、生物免疫系统、航空安全及零信任密码学思想,设计七层防护:(1)基于能力的零信任基础,(2)边界输入验证,(3)意图分类语义防火墙,(4)带安全不变量的对话状态机,(5)行为异常检测,(6)多模型共识验证,(7)分步人工介入升级。构建涵盖8类多轮攻击的正式威胁模型,映射至NIST SP 800-53的14个控制族。理论分析表明,多层防御使攻击成功率降低1-2个数量级。仿真测试覆盖1436个场景(HarmBench 416例,JailbreakBench 200例,XSTest 450例),综合检测率达72.9%(置信区间69.5%-76.0%),经分级响应后有效误报率仅为2.9%,且对多轮渐进式与缓慢漂移攻击实现100%检出。真实基准测试结果(如HarmBench 71.2% vs 作者生成场景76.7%)与对比显示独立评估的重要性。CivicShield填补了人工智能安全、政府合规与实际部署之间的关键空白。

原文摘要 · Abstract (English)

LLM-based chatbots in government services face critical security gaps. Multi-turn adversarial attacks achieve over 90% success against current defenses, and single-layer guardrails are bypassed with similar rates. We present CivicShield, a cross-domain defense-in-depth framework for government-facing AI chatbots. Drawing on network security, formal verification, biological immune systems, aviation safety, and zero-trust cryptography, CivicShield introduces seven defense layers: (1) zero-trust foundation with capability-based access control, (2) perimeter input validation, (3) semantic firewall with intent classification, (4) conversation state machine with safety invariants, (5) behavioral anomaly detection, (6) multi-model consensus verification, and (7) graduated human-in-the-loop escalation. We present a formal threat model covering 8 multi-turn attack families, map the framework to NIST SP 800-53 controls across 14 families, and evaluate using ablation analysis. Theoretical analysis shows layered defenses reduce attack probability by 1-2 orders of magnitude versus single-layer approaches. Simulation against 1,436 scenarios including HarmBench (416), JailbreakBench (200), and XSTest (450) achieves 72.9% combined detection [69.5-76.0% CI] with 2.9% effective false positive rate after graduated response, while maintaining 100% detection of multi-turn crescendo and slow-drift attacks. The honest drop on real benchmarks versus author-generated scenarios (71.2% vs 76.7% on HarmBench, 47.0% vs 70.0% on JailbreakBench) validates independent evaluation importance. CivicShield addresses an open gap at the intersection of AI safety, government compliance, and practical deployment.

AI安全政府AI多层防御对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。