量化AI控制成本,帮开发者权衡安全与开销
Control Tax: The Price of Keeping AI in Check
- 提出'控制税'框架,把安全成本与模型表现挂钩
- 实测主流大模型在对抗攻击下检测漏洞的失败率超40%
- 给出具体成本数据,支持按预算优化监控策略
随着代理型AI在高风险场景中的广泛应用,可靠的监管机制日益重要。本文提出‘控制税’概念,即在AI流程中集成控制措施所产生的操作与财务成本。研究做出三项关键贡献:(1) 构建理论框架,将分类器性能映射为安全保证;(2) 在对抗环境中全面评估前沿语言模型,攻击模型植入隐蔽后门代码,而监测模型尝试识别漏洞;(3) 提供控制协议的实证成本估算,并设计兼顾安全与成本效益的优化监控策略,考虑审计预算等实际约束。该框架使从业者能系统化地将安全承诺与成本关联,推动AI控制在不同部署场景下的可实现性。
原文摘要 · Abstract (English)
The rapid integration of agentic AI into high-stakes real-world applications requires robust oversight mechanisms. The emerging field of AI Control (AIC) aims to provide such an oversight mechanism, but practical adoption depends heavily on implementation overhead. To study this problem better, we introduce the notion of Control tax -- the operational and financial cost of integrating control measures into AI pipelines. Our work makes three key contributions to the field of AIC: (1) we introduce a theoretical framework that quantifies the Control Tax and maps classifier performance to safety assurances; (2) we conduct comprehensive evaluations of state-of-the-art language models in adversarial settings, where attacker models insert subtle backdoors into code while monitoring models attempt to detect these vulnerabilities; and (3) we provide empirical financial cost estimates for control protocols and develop optimized monitoring strategies that balance safety and cost-effectiveness while accounting for practical constraints like auditing budgets. Our framework enables practitioners to make informed decisions by systematically connecting safety guarantees with their costs, advancing AIC through principled economic feasibility assessment across different deployment contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。