arXiv:2607.14006cs.CRcs.AI2026-07

AI系统渗透测试需从破坏基础设施转向检测行为目标违规。

Rethinking Penetration Testing for AI-Enabled Systems: From Resource Compromise to Behavioral Objective Violation

  • 将渗透测试改为基于目标的行为评估,关注系统行为偏离。
  • 识别出10种以上非直接破坏的攻击路径,如提示注入、数据污染等。
  • 适合安全研究人员和AI系统开发者参考,提升对抗性测试能力。

传统渗透测试聚焦于攻击者是否能利用软件、基础设施或配置漏洞实现安全相关破坏,这一范式对人工智能系统依然必要但已不足。在这些系统中,攻击者可通过操纵提示、检索内容、传感器输入、训练数据、记忆、工具或人机交互环路来改变系统行为,而无需直接破坏底层基础设施。本文将针对人工智能系统的渗透测试重新定义为以目标为导向的行为评估。我们定义人工智能系统为学习模型显著影响操作结果的行为系统,并提出‘人工智能渗透’为在明确威胁模型下,可行地引发违反一个或多个操作目标的人工智能主导行为。该定义既保留传统渗透测试,又扩展至提示注入、间接提示注入、数据投毒、传感器操控、检索污染、工具误用及代理对齐失败等路径。我们进一步提出一套测试工作流程:识别操作目标、映射人工智能主导行为、分析攻击影响面、定义行为失效标准、执行场景化测试并报告攻击行为与目标违规之间的证据关联。一个涉及人工智能安全运营中心助手的实例说明了行为影响可能导致渗透,而非基础设施被攻破。整体框架为评估部署中人工智能系统的对抗成功提供了技术基础。

原文摘要 · Abstract (English)

Penetration testing traditionally evaluates whether adversaries can exploit weaknesses in software, infrastructure, configurations, or operational controls to achieve security-relevant compromise. This paradigm remains necessary for AI-enabled systems, but it is no longer sufficient. In such systems, adversaries may influence prompts, retrieved content, sensor inputs, training data, memory, tools, or human-AI interaction loops to alter system behavior without directly compromising the underlying infrastructure. This paper reframes penetration testing for AI-enabled systems as objective-driven behavioral evaluation. We define an AI-enabled system as one in which learned models materially influence behavior affecting operational outcomes, and we define AI-enabled penetration as the feasible induction of AI-governed behavior that violates one or more operational objectives under an explicit threat model. This definition preserves conventional penetration testing while extending it to adversarial pathways such as prompt injection, indirect prompt injection, data poisoning, sensor manipulation, retrieval poisoning, tool misuse, and agentic misalignment. We further propose a testing workflow that identifies operational objectives, maps AI-governed behavior, analyzes adversarial influence surfaces, defines behavioral failure criteria, executes scenario-based tests, and reports evidence linking adversarial action to objective violation. A running example involving an AI-enabled security operations center assistant illustrates how penetration may occur through behavioral influence rather than infrastructure compromise. Together, the definitions, workflow, and example provide a technical framework for evaluating adversarial success in deployed AI-enabled systems.

渗透测试AI安全行为评估对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。