让AI在企业后台任务中合规执行,靠的是带类型检查的计划生成与政策约束。
POLARIS: Typed Planning and Governed Execution for Agentic AI in Back-Office Automation
- 将自动化任务视为带类型检查的有向无环图计划合成。
- 在SROIE数据集上微F1达0.81,异常路由精度达0.95~1.00。
- 适合需可审计、可治理的企业级AI自动化场景。
企业后台流程需要具备可审计性、政策对齐性和操作可预测性的智能体系统,而通用多智能体架构常难以满足。我们提出POLARIS(Policy-Aware LLM Agentic Reasoning for Integrated Systems),一个将自动化视为类型化计划合成与验证执行的受控编排框架。规划器生成结构多样、类型检查的有向无环图(DAGs),基于规则的推理模块选择符合规范的单一计划,执行过程由验证器检查、有限修复循环和编译型策略护栏保障,提前阻断或引导副作用。应用于以文档为中心的金融任务,POLARIS生成决策级成果并保留完整执行轨迹,显著减少人工干预。实证结果显示,在SROIE数据集上微F1为0.81,在可控合成数据集上异常路由精度达0.95至1.00,同时保持审计日志完整。这些评估构成受控智能体AI的初步基准。POLARIS为政策对齐的智能体AI提供了方法论和评估参考。
原文摘要 · Abstract (English)
Enterprise back office workflows require agentic systems that are auditable, policy-aligned, and operationally predictable, capabilities that generic multi-agent setups often fail to deliver. We present POLARIS (Policy-Aware LLM Agentic Reasoning for Integrated Systems), a governed orchestration framework that treats automation as typed plan synthesis and validated execution over LLM agents. A planner proposes structurally diverse, type checked directed acyclic graphs (DAGs), a rubric guided reasoning module selects a single compliant plan, and execution is guarded by validator gated checks, a bounded repair loop, and compiled policy guardrails that block or route side effects before they occur. Applied to document centric finance tasks, POLARIS produces decision grade artifacts and full execution traces while reducing human intervention. Empirically, POLARIS achieves a micro F1 of 0.81 on the SROIE dataset and, on a controlled synthetic suite, achieves 0.95 to 1.00 precision for anomaly routing with preserved audit trails. These evaluations constitute an initial benchmark for governed Agentic AI. POLARIS provides a methodological and benchmark reference for policy-aligned Agentic AI. Keywords Agentic AI, Enterprise Automation, Back-Office Tasks, Benchmarks, Governance, Typed Planning, Evaluation
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。