为AI系统测试环境建立安全评估框架,明确其能力边界与风险防护
AI Sandboxes: A Threat Model, Taxonomy, and Measurement Framework
- 提出基于隔离与监控的AI沙盒架构,划分数字、实体与混合系统类型
- 构建涵盖可信度、可复现性等6维度的量化评估体系,验证于3个真实案例
- 揭示沙盒自身可能被攻击的威胁,适合安全验证与监管合规团队参考
AI系统正越来越多地在包含隔离、仿真、监控、监督和证据留存的受控环境中进行评估。对于物理AI、AIoT及人机协同系统,测试系统可能通过物理过程、网络设备和人类操作感知、决策、执行、通信甚至失效。本文从保障角度出发,将AI沙盒定义为数字AI、具身自主与人机协同部署中用于测试、评估、验证与确认的受控环境。我们形式化了沙盒边界与最弱环节规则,将各维度证据整合为受限部署声明;区分主要沙盒范式;提出包含对保障机制本身攻击在内的网络物理威胁模型;并构建覆盖保真度、可控性、可观测性、封禁性、可复现性及治理文档的测量框架,已在三个真实沙盒案例中实例化。该威胁模型、分类体系与评估框架明确了沙盒可有效验证的内容、可管控的风险以及能支撑的安全、安保与监管证据类型。
原文摘要 · Abstract (English)
AI systems are increasingly evaluated in bounded environments that combine isolation, simulation, instrumentation, supervision, and evidence capture. For physical AI, AIoT, and cyber-physical systems, this shift is not a matter of terminology: the system under test may sense, decide, actuate, communicate, and fail through physical processes, networked devices, and human operators. This article develops an assurance-oriented account of AI sandboxes as controlled environments for testing, evaluation, verification, and validation across digital AI, embodied autonomy, and cyber-physical deployments. We formalize the sandbox boundary and a weakest-link rule for composing per-dimension evidence into a bounded deployment claim; separate major sandbox archetypes; define a cyber-physical threat model that includes attacks on the assurance apparatus itself; and introduce a measurement framework spanning fidelity, controllability, observability, containment, reproducibility, and governance artifacts, instantiated on three worked case studies of real sandboxes. The resulting threat model, taxonomy, and measurement framework clarify what a sandbox can validly test, which risks it can contain, and what forms of evidence it can support for safety, security, and regulatory assurance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。