arXiv:2607.19292cs.CYcs.AI2026-07中稿 · ed被引 4

AI安全不能只看模型输出,更要关注系统如何保持错误可见可控。

The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems

论文配图:The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems
图 1 · 摘自论文原文
  • 提出五层框架,从认知到生态系统诊断隐藏风险
  • 揭示检索依赖、提示注入、奖励劫持等未被重视的隐性威胁
  • 适合关注系统可靠性与治理的开发者和政策制定者

当前AI安全讨论仍过度聚焦于可见的失败,如明显伤害、极端滥用和假想灾难。然而在实际部署中,最具破坏性的故障往往更隐蔽:并非单一输出问题,而是分散在多个组件中,且在流程中被正常化而未被识别为隐患。我们主张,现代AI系统的核心安全挑战不仅是模型是否生成有害内容,更是整个人机技术系统能否维持错误的可见性、可争议性、可控制性和可恢复性。为此,我们提出五层风险诊断框架:(1)认知完整性,确保证据与不确定性被真实呈现以支持合理信赖;(2)控制完整性,保障权限与行动边界在攻击和优化下仍稳健;(3)时间完整性,确保安全在会话延续、记忆更新和部署漂移中持续有效;(4)组织完整性,保证机构具备审计、追责和干预能力;(5)生态完整性,确保AI不侵蚀未来监管所依赖的信息环境。在此基础上,识别出一系列被忽视的风险模式,包括过度依赖、检索中的不确定性与合法性洗白、提示注入、奖励劫持、记忆污染、评估欺骗、虚构人类监督、合成证据污染及模型坍塌。最后,提出设计与治理建议,并呼吁将AI安全研究从模型中心转向社会技术系统的可靠性。

原文摘要 · Abstract (English)

Current AI safety discourse still focuses disproportionately on visible failures, including obvious harms, dramatic misuse, and hypothetical catastrophic scenarios. That focus is incomplete. In deployed systems, many of the most consequential failures are quieter: plausible rather than spectacular, distributed across components rather than localized in a single output, and normalized by workflows before they are recognized as hazards. We argue that a central safety challenge in modern AI systems is increasingly not only whether a model emits a harmful response, but whether the broader socio-technical system preserves the conditions under which errors remain visible, contestable, containable, and recoverable. We propose a five-layer framework for diagnosing these hidden risks: (1) epistemic integrity, concerning whether evidence and uncertainty are represented honestly enough to support calibrated reliance; (2) control integrity, concerning whether authority, permissions, and action boundaries remain robust under attack and optimization; (3) temporal integrity, concerning whether safety holds across sessions, memory updates, and deployment drift; (4) organizational integrity, concerning whether institutions retain the capacity to audit, assign responsibility, and intervene effectively; and (5) ecosystem integrity, concerning whether AI systems preserve rather than erode the information environment on which future oversight depends. Across these layers, we identify under-recognized risk patterns, including overreliance, uncertainty and legitimacy laundering in retrieval, prompt injection, reward hacking, memory poisoning, evaluation deception, fictional human oversight, synthetic evidence pollution, and model collapse. We conclude with design and governance recommendations and a research agenda for shifting AI safety from model-centric evaluation toward socio-technical reliability.

AI安全系统风险治理框架隐性威胁

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。