arXiv:2607.21268cs.MAcs.AI2026-07

用可控的多人协作机制,让AI更可靠地构建经济学理论。

pAI-Econ-claude: A Gated Human-in-the-Loop Multi-Agent Architecture for AI-Assisted Economic Theory Development

论文配图:pAI-Econ-claude: A Gated Human-in-the-Loop Multi-Agent Architecture for AI-Assisted Economic Theory Development
图 1 · 摘自论文原文
  • 通过可检查的中间记录和分层审查门控,实现AI与人类协同决策。
  • 错误严重度下降26.6%,有用性提升19.2%,关键环节纠错效果显著。
  • 适合需要高可信度经济推演的研究者,尤其关注可审计性而非完全自动化。

在经济学等社会科学任务中,基于大模型的智能体生成结果缺乏低成本、完整且机器可读的正确性信号,导致多智能体系统面临独特可靠性挑战:当无任何组件能认证最终结果时,如何组织生成、批判、协调与人类判断?我们提出pAI-Econ-claude,一种带门控的、人机协同的多智能体架构,用于辅助经济理论发展。智能体通过可检查的共享工作空间进行协作;专用门控识别特定失败模式并建议回退,但不认证正确性;人类检查点保留对高成本不可逆决策的最终控制权。在五个匹配的经济理论任务上,与无门控基线对比评估。两位盲评者对所有五组配对排名达成一致,偏好带门控架构四次,基线一次。平均错误严重度从1.58降至1.16,整体有用性从2.60升至3.10。最大收益来自现实检验否决了错误的市场结构假设,以及证明审查触发了错误福利主张的修正。负面案例显示,框架也可能过度压缩重要经济机制。结果支持有限结论:门控监督提升了AI辅助经济理论的可审计性,但无法替代正式验证;不可逆人类判断的分配比纯代理自主性更具设计意义。工作流已公开于https://github.com/maxwell2732/pAI-Econ-claude。

原文摘要 · Abstract (English)

In many social-science research tasks, such as economics, LLM-based agents must produce outputs for which no cheap, task-complete, machine-readable correctness signal exists. This creates a distinctive reliability problem for multi-agent systems: how should generation, critique, coordination, and human judgment be organized when no component can certify the final result? We address this problem through pAI-Econ-claude, a gated, human-in-the-loop multi-agent architecture for AI-assisted economic theory development. Agents coordinate through a shared workspace of inspectable intermediate records; specialized gates diagnose targeted failure modes and recommend loopbacks without certifying correctness; and human checkpoints retain authority over decisions that are costly to reverse. We evaluate the architecture on five matched economic-theory tasks against an ungated baseline. Two evaluators blinded to configuration agreed on all five pairwise rankings, preferring the gated architecture in four tasks and the baseline in one. Mean failure severity fell from 1.58 to 1.16, while overall usefulness rose from 2.60 to 3.10. The largest observed gain occurred when a reality check rejected a false market-structure premise and a proof review prompted revision of a false welfare claim. The negative case shows that scaffolding can also compress an economically important mechanism too aggressively. The results support a bounded claim: gated oversight improves the auditability of AI-assisted economic theory without substituting for formal verification, and the allocation of irreversible human judgment is a more informative design variable than pure agent autonomy. The workflow is publicly available at https://github.com/maxwell2732/pAI-Econ-claude.

多智能体人机协同经济建模可审计性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。