arXiv:2607.29405cs.AIcs.MA2026-07

验证智能体系统需关注动态行为而非单一组件性能。

Beyond Component Testing: Validating Agentic AI Systems

论文配图:Beyond Component Testing: Validating Agentic AI Systems
图 1 · 摘自论文原文
  • 构建五维评估框架:行为、安全、时间、合规与多智能体交互
  • 发现时间有效性与运行时证据维护仍是薄弱环节
  • 适合安全关键领域研究者与系统设计者参考

智能体系统通过多步决策轨迹实现规划、工具使用、记忆、交互与适应,其行为跨越了传统组件测试与单次输入输出评估的范畴。本文综述257篇文献,涵盖智能体评估、软件保障、网络物理系统、运行时监控与监管指引,提出包含行为、安全、时间、合规与多智能体五个维度的评估框架,并据此映射现有方法、揭示覆盖盲区。分析显示行为评估相对成熟,但时间有效性、运行时证据留存、合规可读性及开放环境多智能体系统保障仍待发展。三个跨领域案例(医疗护理、工业运营、智慧出行)展示五维维度在安全关键场景中的实际体现,基于文献中记录的失效模式。论文最后提出以生命周期为导向的研究议程,聚焦有限自主规范、对抗性轨迹生成、运行时监控与可审计证据结构。核心观点:可信部署依赖对上下文中的行为轨迹进行验证,而非孤立评估组件。

原文摘要 · Abstract (English)

Agentic AI systems act through multi-step trajectories that combine planning, tool use, memory, interaction, and adaptation. This behavior stretches validation practice beyond component testing and one-shot input--output evaluation, because acceptable system behavior now depends on how decisions unfold over time and under changing environmental conditions. This survey synthesizes 257 papers spanning agent evaluation, software assurance, cyber-physical systems, runtime monitoring, and regulatory guidance in order to characterize the validation problem for agentic systems. The review is organized around a five-dimension taxonomy covering behavioral, safety, temporal, regulatory, and multi-agent concerns, and uses that taxonomy to map current approaches and expose recurrent coverage gaps. The analysis shows that behavioral evaluation is comparatively mature, while temporal validity, runtime evidence maintenance, regulatory legibility, and open-ended multi-agent systems assurance remain under-developed. Three cross-domain case studies (medical care, industrial operations, smart-mobility systems) provide operational illustrations of how the five taxonomy dimensions recur in safety-critical settings, grounded in the failure patterns documented in the reviewed literature. The paper concludes with a lifecycle-oriented research agenda centered on bounded-autonomy specifications, adversarial trajectory generation, runtime monitoring, and audit-ready evidence structures. The central claim is that trustworthy deployment of agentic AI depends on validating trajectories in context rather than assessing isolated components alone.

智能体系统验证评估安全关键生命周期

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。