评估AI代理安全不能只看单点表现,要关注人机协同的可靠性。
Beyond single-channel agentic benchmarking
- 提出多通道评估思路,强调人机系统中错误模式的独立性
- 实验显示不完美的AI可作为冗余审计层,降低人为失误风险
- 适合关注人机协作安全的开发者与评估者参考
当前针对智能体人工智能的安全评估多基于单一任务准确率阈值,将自主系统视为单一故障点,这与安全关键工程中通过冗余、错误模式多样性及系统联合可靠性来缓解风险的原则相悖。本文指出,在人机协同环境中孤立评估AI代理会系统性误判其安全性。以近期实验室安全基准为案例,表明即使性能不完善的AI也能通过作为冗余审计层,有效应对人类常见的失效模式,如警觉性下降、无意盲视和偏差正常化。该研究将评估重点从代理绝对准确性转向人机二元系统的可靠性,尤其强调错误模式的非相关性是降低风险的核心因素。这一视角使AI评估更符合其他安全关键领域的实践,为生态有效的安全测评提供路径。
原文摘要 · Abstract (English)
Contemporary benchmarks for agentic artificial intelligence (AI) frequently evaluate safety through isolated task-level accuracy thresholds, implicitly treating autonomous systems as single points of failure. This single-channel paradigm diverges from established principles in safety-critical engineering, where risk mitigation is achieved through redundancy, diversity of error modes, and joint system reliability. This paper argues that evaluating AI agents in isolation systematically mischaracterizes their operational safety when deployed within human-in-the-loop environments. Using a recent laboratory safety benchmark as a case study demonstrates that even imperfect AI systems can nonetheless provide substantial safety utility by functioning as redundant audit layers against well-documented sources of human failure, including vigilance decrement, inattentional blindness, and normalization of deviance. This perspective reframes agentic safety evaluation around the reliability of the human-AI dyad rather than absolute agent accuracy, with a particular emphasis on uncorrelated error modes as the primary determinant of risk reduction. Such a shift aligns AI benchmarking with established practices in other safety-critical domains and offers a path toward more ecologically valid safety assessments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。