arXiv:2607.18548cs.AIcs.MA2026-07被引 1

为关键系统中的智能体AI构建可验证、可审计的信任框架

Engineering Trustworthy Agentic AI for Critical Systems

论文配图:Engineering Trustworthy Agentic AI for Critical Systems
图 1 · 摘自论文原文
  • 从安全、鲁棒性、可解释性等五个维度构建信任模型
  • 覆盖感知到审计全流程,提出可量化的信任评估机制
  • 跨电力、自动驾驶、高性能计算等四领域提炼通用设计模式

具备自主感知、规划、工具使用和多步行动能力的智能体人工智能系统正被越来越多地应用于关键工程领域,其决策可能带来物理、运营或经济后果。本文将可信性——即智能体行为在实际工程约束下是否可验证、可审计、可信赖——作为首要工程属性,而非仅以任务能力评价智能体AI。研究构建了一个涵盖安全与约束满足、鲁棒性与可靠性、透明性与可解释性、问责性与可审计性、隐私与安全五个核心维度的可信性模型,并映射至从感知到审计的全流程保障工作流。在此基础上,系统梳理了智能体架构、威胁、具体信任机制及量化指标,支持直接应用于智能体系统开发与评估。进一步在电力系统、自动驾驶/机器人/无人机、高性能计算、通信网络四个受约束的工程领域展开分析,识别出共性设计模式、共享失效模式及领域特异性差距。综合来看,智能体AI的可信性是一个统一问题,本文提出一条通往可复用、跨领域的保障框架的路径,类似成熟安全关键工程领域的分级认证制度。

原文摘要 · Abstract (English)

Agentic artificial intelligence systems, capable of autonomous perception, planning, tool use, and multi-step action, are increasingly proposed for critical engineering domains where decisions carry physical, operational, or economic consequences. This survey addresses a gap in current literature by treating trustworthiness, whether agentic behavior can be verified, audited, and trusted under the constraints that engineering practice actually requires, as a first-class engineering property, rather than evaluating agentic AI by task capability alone. The study adopts a trustworthiness model organized around five cross-cutting dimensions: safety and constraint satisfaction; robustness and reliability; transparency and interpretability; accountability and auditability; and privacy and security. This is mapped onto an agentic assurance workflow spanning perception through audit. Building on this foundation, agentic systems architectures, threats, concrete trust mechanisms, and quantitative metrics are surveyed for direct application in agentic systems development and evaluation. These principles are then examined across four constraint-bound engineering domains: power systems, autonomous vehicles/robotics/UAVs, high-performance computing, and communication networks, identifying recurring design patterns, shared failure modes, and domain-specific gaps. Synthesizing across those domains, agentic AI trustworthiness is shown to be a single problem, with a path outlined toward a reusable, cross-domain assurance framework analogous to the graded certification regimes used by mature safety-critical engineering fields.

智能体AI可信性工程系统安全保障

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。