用公共品稳定菜单问题测试AI经济学研究流程,发现人类直觉和多轮交互提升效果。
Stable Menus of Public Goods: AI-Enabled Progress
- 用人类直觉引导提示,让大模型产生更好判断
- 多轮交互在鼓励大胆探索时更有效
- 大模型略逊于一年级博士生,但差距小
以EC 2025论文中的开放问题为实验基准,我们评估不同AI助力经济计算研究的工作流效果。重点考察三个问题:在提示中加入人类直觉是否有帮助?自动化多轮交互是否有效?大语言模型(LLM)是否优于一年级博士生?针对前两个问题,我们发现:(1)以人类直觉为提示能促使LLM展现出更优的“品味”;(2)当流程鼓励“雄心勃勃”的步骤时,多轮交互更有优势。针对第三个问题,基于资深作者未合作前撰写的未发表手稿,我们对比了LLM与一年级博士生的表现,结果表明LLM略逊一筹,但差距有限。
原文摘要 · Abstract (English)
Using an open problem from the EC 2025 paper "Stable Menus of Public Goods" as a testbed, we conduct experiments to understand the effectiveness of different AI-for-EconCS research workflows. Specifically, we study three questions: Does providing human intuition in the prompt help? Does automated multi-turn interaction help? And, does an LLM outperform a first-year PhD student? Regarding the first two questions, we provide evidence for the following workflow suggestions: (1) prompting with human intuition can encourage the LLM to have better "taste", (2) multi-turn workflows help when the pipeline encourages "ambitious" steps. Regarding the third question, using an unpublished manuscript written by the paper's senior authors prior to collaborating with the first-year PhD student, we compare the effectiveness of the LLM with that of the first-year PhD student, and find that the LLM is slightly less effective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。