arXiv:2609.05531cs.CLcs.AI2026-09

AI治理需匹配模型能力,好框架能显著提升医疗任务成功率。

When Agent Governance Helps

论文配图:When Agent Governance Helps
图 1 · 摘自论文原文
  • 设计可运行的GAMPO框架,将自治与监管结合,实现目标自生成。
  • 在医疗长周期任务中,治理效果依赖模型冗余能力,前沿模型收益更高。
  • 按具体案例定制验收标准比统一流程更有效,适合高阶AI应用者。

现有研究未明确如何设计和评估在约束下自主追求目标的智能体组织。本文分两步回应:首先,基于321份文献的定性证据整合,提出可运行的受控自治多智能体产品组织(GAMPO)框架,融合代理、敏捷、平台与治理理论;其次,在CHI-Bench医疗长周期基准上测试其提示层实例,覆盖开源与前沿模型。结果发现:治理效益由模型冗余能力决定,且具领域与模型特异性。在容量受限的开源模型中,完整流程无可靠增益;而仅添加“验证你的写入”一句,任务通过率从2/20升至4/20。在前沿模型上,相同架构使预授权成功率从24%提升至40%,但另一模型无增益,根源在于稳定推荐覆写倾向。进一步优化:用仅依据案例政策与公开标准的盲答案定义完成条件替代通用流程,结合五次自洽优选,使预授权达84%(单次68%),利用率管理达44%,但护理管理受内容质量瓶颈制约。贡献为命名可审计的框架与能力感知证据,表明治理应适配冗余容量,且在前沿阶段案例驱动规范优于统一流程。研究具探索性:部分实例化、小样本(n=5–25)、单次试验。

原文摘要 · Abstract (English)

No specification says how a governed autotelic AI agent organization, where agents pursue self-generated goals inside guardrails, should be designed and evaluated. We answer in two parts. First, we synthesize the Governed Autotelic Multi-Agent Product Organization (GAMPO) framework from a document-based qualitative evidence synthesis of 321 sources, integrating agency, agile, platform, and governance theory into a runnable specification. Second, we probe a prompt-layer instantiation of GAMPO on CHI-Bench, a long-horizon healthcare benchmark, across open and frontier models. The result is a boundary condition: governance benefit is gated by a model's spare capacity and is domain- and model-specific. On capacity-constrained open models the full procedure yields no reliable benefit, whereas a single "verify your writes" sentence doubles task success (pass@1 2/20 to 4/20). At the frontier the same scaffold lifts prior-authorization 24% to 40% but nets zero on another model, a gap traced to a stable recommendation-override disposition. A second result refines the first: replacing the generic procedure with an answer-blind, per-task definition-of-done, keyed only to the case's own policy and published standards, never the hidden key, raises prior-authorization to 84% under best-of-five self-consistency (68% single-attempt, confirmed by a held-out board) and utilization-management to 44%, while care-management meets a subjective content-quality wall. The contribution is a named, auditable framework and capability-gated evidence that governance should be sized to spare capacity, and that at the frontier a case-grounded specification beats a uniform procedure. Findings are exploratory: partial instantiation, small per-cell samples (n = 5-25), and single trials.

AI治理多智能体医疗AI框架设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。