为自主智能系统设计双层可靠性监控框架,提升高风险场景下的安全性。
Perspectives on a Reliability Monitoring Framework for Agentic AI Systems
- 提出双层监控框架:检测异常输入与揭示内部决策过程。
- 实现运行时可靠性评估,支持人工干预以降低意外行为风险。
- 适用于医疗、工业等对可靠性要求高的场景。
自主智能系统在多种应用中具有提供更高效帮助的潜力,其可自主达成目标且外部干预较少。然而,其可靠性不足使其难以应用于医疗或流程工业等高风险领域。不可靠系统在运行中可能产生意外行为,需采取缓解措施。本文基于自主智能系统的特性,分析了其运行中的主要可靠性挑战,并与传统人工智能系统进行关联,提出了一个固有的可靠性挑战。作为主要贡献,我们提出一种两层可靠性监控框架:第一层为分布外检测,用于识别新颖输入;第二层为人工智能透明度层,用于揭示内部运作机制。该框架为人机协同决策提供支持,使操作员能够判断输出是否可能不可靠并及时干预。该框架为开发缓解技术奠定了基础,有助于降低运行中因可靠性不确定带来的风险。
原文摘要 · Abstract (English)
The implementation of agentic AI systems has the potential of providing more helpful AI systems in a variety of applications. These systems work autonomously towards a defined goal with reduced external control. Despite their potential, one of their flaws is the insufficient reliability which makes them especially unsuitable for high-risk domains such as healthcare or process industry. Unreliable systems pose a risk in terms of unexpected behavior during operation and mitigation techniques are needed. In this work, we derive the main reliability challenges of agentic AI systems during operation based on their characteristics. We draw the connection to traditional AI systems and formulate a fundamental reliability challenge during operation which is inherent to traditional and agentic AI systems. As our main contribution, we propose a two-layered reliability monitoring framework for agentic AI systems which consists of a out-of-distribution detection layer for novel inputs and AI transparency layer to reveal internal operations. This two-layered monitoring approach gives a human operator the decision support which is needed to decide whether an output is potential unreliable or not and intervene. This framework provides a foundation for developing mitigation techniques to reduce risk stemming from uncertain reliability during operation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。