arXiv:2604.25684cs.AI2026-04

让AI像人一样思考再行动,实现自主决策中的安全自控。

Think Before You Act -- A Neurocognitive Governance Model for Autonomous AI Agents

论文配图:Think Before You Act -- A Neurocognitive Governance Model for Autonomous AI Agents
图 1 · 摘自论文原文
  • 构建前置行动治理推理循环,嵌入四层规则审查机制。
  • 在供应链场景中达成95%合规准确率,零误报人工干预。
  • 适合高风险领域如医疗、金融的自主AI系统参考应用。

自主AI代理在企业、医疗及关键安全场景的快速部署,暴露出根本性的治理缺口。现有方法如运行时约束、训练对齐和事后审计,将治理视为外部强加的限制,而非内化的行为准则,导致代理易产生不安全且不可逆的行为。本文借鉴人类自我治理机制:行动前通过执行功能、抑制控制和内化规则进行审慎评估,判断行为是否合规、需修改或需上报。提出神经认知治理框架,将人类自控过程形式化映射至大语言模型驱动的代理推理,建立人脑与模型认知核心之间的结构类比。引入预行动治理推理循环(PAGRL),要求代理在每次重要操作前,调用全局、工作流特定、代理特定和情境四层规则集进行审查,模拟企业组织中跨层级合规架构。在生产级零售供应链流程中验证,该框架实现95%合规准确率,零误报人工干预,证明将治理嵌入推理过程,可带来更一致、可解释、可审计的合规性,优于外部强制。本工作为自主AI提供一种原则性基础:不是因为规则被强加,而是因为审慎思考已融入其认知逻辑。

原文摘要 · Abstract (English)

The rapid deployment of autonomous AI agents across enterprise, healthcare, and safety-critical environments has created a fundamental governance gap. Existing approaches, runtime guardrails, training-time alignment, and post-hoc auditing treat governance as an external constraint rather than an internalized behavioral principle, leaving agents vulnerable to unsafe and irreversible actions. We address this gap by drawing on how humans self-govern naturally: before acting, humans engage deliberate cognitive processes grounded in executive function, inhibitory control, and internalized organizational rules to evaluate whether an intended action is permissible, requires modification, or demands escalation. This paper proposes a neurocognitive governance framework that formally maps this human self-governance process to LLM-driven agent reasoning, establishing a structural parallel between the human brain and the large language model as the cognitive core of an agent. We formalize a Pre-Action Governance Reasoning Loop (PAGRL) in which agents consult a four-layer governance rule set: global, workflow-specific, agent-specific, and situational before every consequential action, mirroring how human organizations structure compliance hierarchies across enterprise, department, and role levels. Implemented on a production-grade retail supply chain workflow, the framework achieves 95% compliance accuracy and zero false escalations to human oversight, demonstrating that embedding governance into agent reasoning produces more consistent, explainable, and auditable compliance than external enforcement. This work offers a principled foundation for autonomous AI agents that govern themselves the way humans do: not because rules are imposed upon them, but because deliberation is embedded in how they think.

AI治理自主代理大模型安全决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。