为大模型企业系统设计可落地的风险评估框架
AI Assurance: A Comprehensive Testing Strategy for Enterprise AI Systems

- 以持续降险替代传统正确性验证,重构AI测试逻辑
- 提出五层保障金字塔,覆盖从开发到治理全流程
- 适合关注AI系统可靠性的工程管理者与实践者
企业级AI系统基于大语言模型、检索管道和自主代理构建,带来传统软件质量保障无法应对的新风险。这些系统具有概率性、上下文敏感性和涌现性:无法像经典程序那样被严格验证正确,只能通过不断积累信心来评估。本文提出一套围绕三大原则的综合保障策略:第一,AI测试应聚焦持续降低风险而非追求绝对正确;第二,评估必须作为与开发并重的核心工程学科;第三,AI保障失败可能引发与传统确定性软件截然不同的组织影响。论文引入结构化AI故障分类体系,提出改进的五层AI保障金字塔,并提供评估驱动开发、RAG系统测试、模型生命周期管理及治理的实操指南。目标是为工程领导者和实践者提供兼具哲学基础与可操作性的策略。
原文摘要 · Abstract (English)
Enterprise AI systems, built on large language models, retrieval pipelines and autonomous agents, introduce a class of risks that traditional software quality assurance was never designed to address. These systems are probabilistic, context-sensitive and emergent: they cannot be verified to be correct in the classical sense, but only evaluated with increasing confidence. This paper presents a comprehensive assurance strategy for enterprise AI systems built around three key principles: first, that AI testing should focus on continuous risk reduction rather than strict correctness verification; second, that evaluation must be treated as a core engineering discipline alongside development; and third, that failures in AI assurance can lead to organizational impacts that are fundamentally different from those seen in traditional deterministic software systems. We introduce a structured AI Failure Taxonomy, propose a revised five-layer AI Assurance Pyramid and provide operational guidance on evaluation-driven development, RAG system testing, model lifecycle management and governance. The goal is to equip engineering leaders and practitioners with a strategy that is both philosophically grounded and operationally deployable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。