用多智能体流程图提升AI合规工作流的可靠性与效率
Constrained Process Maps for Multi-Agent Generative AI Workflows
- 将多步合规任务建模为带有明确角色分工的马尔可夫决策过程
- 相比单智能体,准确率提升19%,人工审核量减少85倍
- 适合需要可解释性与人类监督的高风险AI应用开发者
基于大语言模型的智能体在合规、尽职调查等受监管场景中日益广泛应用。然而,多数架构依赖单一智能体的提示工程,难以观察或比较模型在不确定性处理和跨阶段协作中的表现,也缺乏对人类干预的清晰机制。本文提出一种多智能体系统,形式化为具有有向无环结构的有限时域马尔可夫决策过程(MDP),每个智能体对应特定角色或决策阶段(如内容、业务或法律审查),预定义的转移规则表示任务升级或完成。通过蒙特卡洛估计量化个体智能体的信念不确定性,系统级不确定性则由终止于自动标注状态或人工审查状态来捕捉。以自伤检测的AI安全评估为例,该系统实现性能提升:相比单智能体基线,准确率最高提高19%,所需人工审核减少最多85倍,部分配置下处理时间也缩短。
原文摘要 · Abstract (English)
Large language model (LLM)-based agents are increasingly used to perform complex, multi-step workflows in regulated settings such as compliance and due diligence. However, many agentic architectures rely primarily on prompt engineering of a single agent, making it difficult to observe or compare how models handle uncertainty and coordination across interconnected decision stages and with human oversight. We introduce a multi-agent system formalized as a finite-horizon Markov Decision Process (MDP) with a directed acyclic structure. Each agent corresponds to a specific role or decision stage (e.g., content, business, or legal review in a compliance workflow), with predefined transitions representing task escalation or completion. Epistemic uncertainty is quantified at the agent level using Monte Carlo estimation, while system-level uncertainty is captured by the MDP's termination in either an automated labeled state or a human-review state. We illustrate the approach through a case study in AI safety evaluation for self-harm detection, implemented as a multi-agent compliance system. Results demonstrate improvements over a single-agent baseline, including up to a 19\% increase in accuracy, up to an 85x reduction in required human review, and, in some configurations, reduced processing time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。