用提示词经济实现大规模跨学科自动研究,机器做执行,人类掌方向。
Agon: An Autonomous Large-Scale Omnidisciplinary Research System Built on Prompt Economy

- 基于提示词经济构建自动化研究系统,仅用小主题启动,无需人工写代码。
- 444轮循环中实现跨领域研究,发现新型失败模式并分类管理。
- 提出可修复性、可见性等四维度失败分类法,明确机器与人分工边界。
大型语言模型正使研究产出规模化,瓶颈从生成成果转向判断结论。我们提出 extsc{Agon},一个研究协调系统,能在流程内验证可检查的内容,其余判断交由人类科学家完成。 extsc{Agon} 遵循六项设计原则:提示词经济、面向未来、极简提示、跨学科、大规模并行和零代码。我们在多个领域运行了 444 次提示词经济循环,仅使用小型起始主题且未编写任何人工实验代码。这些部署展示了系统的可扩展性,同时揭示了新类别的失败现象。我们将这些失败按严重性、可修复性、可见性和能力归属进行分类,区分出系统可识别并修复的失败与需人类判断的部分。结果表明, extsc{Agon} 正推动研究进入新范式:机器承担规模任务,人类主导方向把控。
原文摘要 · Abstract (English)
Large language models are making research production scalable, shifting the bottleneck from producing artifacts to judging claims. We present \textsc{Agon}, a research orchestrator that validates what can be checked inside the workflow and leaves the remaining judgments to human scientists. \textsc{Agon} is built on six design principles: Prompt Economy, Future-Facing, Minimal Prompts, OmniDisciplinary, Massive Parallelism, and Zero-Code. We ran \textsc{Agon} across domains for 444 iterations of Prompt Economy loops, using only small starting topics and no human-written experimental code. These deployments demonstrate scalability while exposing new classes of failure. We organize these failures into a taxonomy along severity, fixability, visibility, and capability locus. The taxonomy separates failures the loops can see and fix from those that require human judgment. Together, these results show that \textsc{Agon} is pushing research toward a new paradigm: machine scales, human steers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。