用知识图谱生成测试场景,提前验证企业AI代理的合规与安全。
Toward Pre-Deployment Assurance for Enterprise AI Agents: Ontology-Grounded Simulation and Trust Certification

- 基于知识图谱自动生成监管、运营和对抗性测试场景。
- 在4个行业10个监管单元中覆盖48.3%的法规要求,优于传统方法。
- 生成可机器验证的信任证书,适合金融医疗等强监管领域使用。
企业人工智能代理的部署前验证仍存在关键空白,现有后置监控与人工干预无法提供充分保障。本文提出首个融合三要素的本体驱动验证框架:定义代理操作边界以涵盖权限、领域约束、安全属性、治理规则与自主等级;构建本体到场景的自动生成功能,生成监管、运营及对抗性测试案例;产出可机器验证的分级信任证书。在美越两国四个受监管行业(金融科技、银行、保险、医疗)的试点中,覆盖125项原始法规要求,注入25个故障,生成1800个测试场景。本体生成法在法规覆盖率上显著优于主流人物角色基线(48.3% vs 33.1%;校正后p_c=0.0006),领域特异性达4.77/5.0(p=2e-6);其对普通提示与检索增强提示的优势在贝叶斯校正后不显著。跨三大LLM家族(Claude Sonnet 4、Qwen 2.5 72B、Gemma 4 26B;共5400场景)的交叉验证复现了本体优于人物角色的模式。该框架为企事业单位提供可复现、法规驱动的部署前保障路径,补足运行时治理,形成可审计的部署准入机制。
原文摘要 · Abstract (English)
Pre-deployment verification of enterprise artificial intelligence (AI) agents remains a critical gap between large language model (LLM) capability benchmarking and production deployment. Post-deployment monitoring, human-in-the-loop controls, and prompt-level guardrails offer limited assurance once an agent is operating in production. We present an ontology-grounded verification framework -- to our knowledge the first to combine three components: an Agent Operational Envelope formalizing the certification space across permissions, domain constraints, safety properties, governance rules, and autonomy levels; an ontology-to-scenario generation pipeline that derives regulatory, operational, and adversarial test scenarios automatically; and a machine-verifiable Trust Certificate with graduated deployment verdicts. A controlled pilot across four regulated industries (Fintech, Banking, Insurance, Healthcare), instantiated as five industry-by-regulatory-regime cells across the United States and Vietnam (where Vietnam's 2025 AI Law makes such verification legally mandated for financial services), generated 1,800 scenarios evaluated against 125 primary-source regulatory requirements and 25 injected faults. Ontology-grounded generation significantly outperformed the dominant persona-based baseline on regulatory coverage (48.3% versus 33.1%; corrected p_c = .0006) and attained the highest domain specificity (4.77/5.0; p = 2e-6); transparently, its advantage over plain and retrieval-augmented prompting did not survive Bonferroni correction. Cross-validation across three LLM families (Claude Sonnet 4, Qwen 2.5 72B, Gemma 4 26B; 5,400 total scenarios) replicated the persona-versus-ontology pattern. The framework offers a reproducible, regulation-grounded route to pre-deployment assurance for enterprise AI agents, complementing runtime governance with an auditable deployment gate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。