PARC能自主执行复杂计算任务并自我纠错,无需人工干预。
PARC: An Autonomous Self-Reflective Coding Agent for Robust Execution of Long-Horizon Tasks
- 分层多智能体架构,自带自我评估与反馈机制
- 完成数十个并行模拟任务,每项耗时约43小时
- 从自然语言指令出发,生成可竞争人类基准的分析方案
我们提出PARC,一个用于自主、鲁棒执行长时程计算任务的编程智能体。PARC基于分层多智能体架构,包含任务规划、执行以及从独立上下文评估自身行为与结果并提供反馈的机制,即自我评估与自我反馈。该设计使PARC能够检测并纠正高层战略错误,持续推进任务而无需人工干预。我们在计算科学与数据科学任务中评估PARC:在材料科学中,它自主复现了锂离子传导与合金偏析研究的关键结果;具体而言,协调数十个并行仿真任务,每个任务约需43小时计算,实现端到端的调度、监控与错误修正。在基于Kaggle的实验中,从极简自然语言指令出发,PARC完成数据分析与搜索策略实施,生成解决方案可媲美人工工程基线。这些结果表明,将分层多智能体系统与自我评估、自我反馈相结合,有望实现具备独立开展大规模科学与分析工作的AI系统。
原文摘要 · Abstract (English)
We introduce PARC, a coding agent for the autonomous and robust execution of long-horizon computational tasks. PARC is built on a hierarchical multi-agent architecture incorporating task planning, execution, and a mechanism that evaluates its own actions and their outcomes from an independent context and provides feedback, namely self-assessment and self-feedback. This design enables PARC to detect and correct high-level strategic errors and sustain progress without human intervention. We evaluate PARC across computational science and data science tasks. In materials science, it autonomously reproduces key results from studies on lithium-ion conduction and alloy segregation. In particular, it coordinates dozens of parallel simulation tasks, each requiring roughly 43 hours of computation, managing orchestration, monitoring, and error correction end-to-end. In Kaggle-based experiments, starting from minimal natural-language instructions, PARC conducts data analysis and implements search strategies, producing solutions competitive with human-engineered baselines. These results highlight the potential of integrating a hierarchical multi-agent system with self-assessment and self-feedback to enable AI systems capable of independent, large-scale scientific and analytical work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。