arXiv:2512.22154cs.CRcs.AI2025-12被引 6

研究前沿AI部署中监控系统的实际挑战与应对策略

Practical challenges of control monitoring in frontier AI deployments

  • 提出同步、半同步、异步三种监控模式,权衡延迟与安全
  • 发现监督、延迟、恢复是三大核心挑战,需在真实场景中综合考虑
  • 适用于关注AI安全与可靠性研究的从业者和安全架构设计者

自动化控制监控在监督不可完全信任的高能力AI代理中可能发挥重要作用。以往研究多在简化环境中探索监控机制,但将监控扩展至真实部署时引入了新动态:多个代理并行运行、监督存在不可忽略的延迟、代理间存在渐进式攻击,且仅凭单次有害行为难以识别隐藏意图的代理。本文分析了应对这些挑战的设计选择,聚焦三种具有不同延迟-安全权衡的监控形式:同步、半同步和异步监控。提出一个高层级的安全论证框架,用以理解与比较不同监控协议。分析揭示了三大关键挑战——监督、延迟与恢复,并通过四个未来AI部署案例进行探讨。

原文摘要 · Abstract (English)

Automated control monitors could play an important role in overseeing highly capable AI agents that we do not fully trust. Prior work has explored control monitoring in simplified settings, but scaling monitoring to real-world deployments introduces additional dynamics: parallel agent instances, non-negligible oversight latency, incremental attacks between agent instances, and the difficulty of identifying scheming agents based on individual harmful actions. In this paper, we analyse design choices to address these challenges, focusing on three forms of monitoring with different latency-safety trade-offs: synchronous, semi-synchronous, and asynchronous monitoring. We introduce a high-level safety case sketch as a tool for understanding and comparing these monitoring protocols. Our analysis identifies three challenges -- oversight, latency, and recovery -- and explores them in four case studies of possible future AI deployments.

AI安全监控系统延迟可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。