评测大模型代理在机器学习任务中的创造力,发现其创新但难落地。
Can LLM Agents Discover? Evaluating Creativity on ML Engineering Tasks
- 用多轮交互任务评估大模型代理的创意,从心理、历史和实用性三维度出发。
- 模型比人类更会创新,但实际性能反而更低,说明创意未转化为有效成果。
- 开发自动化评判系统,可大规模可靠评估模型的创意水平,适合算法研究者参考。
近期人工智能系统宣称能实现自主科学发现,包括设计算法并撰写论文,但其是否具备创造力——即产生既新颖又实用的解决方案的能力——仍不明确。本文提出一个基于机器学习工程任务的评估框架,通过三个维度衡量多轮对话的大模型研究代理的创造力:心理新颖性(P-Creativity,相对于代理自身先前解法的新颖性)、历史新颖性(H-Creativity,相对于人类已有解法的新颖性)以及实用性(任务表现)。在10个来自MLE-Bench的类似Kaggle的机器学习任务上,对AIDE与AIRA-Dojo两个代理框架进行评估,构建了基于大模型作为裁判的自动化评估流水线,并验证其与人工创造力判断高度相关,实现了大规模、可靠的P-Creativity评估。分析代理行为轨迹发现:(1) 所有代理在从探索转向利用的过程中,心理新颖性持续下降;(2) 大模型表现出高于获奖人类的H-Creativity,但任务性能更低。结果表明,当前代理虽能探索解空间的新区域,却缺乏将新颖性转化为更高性能的能力。
原文摘要 · Abstract (English)
Recent AI systems promise autonomous scientific discovery, claiming to discover algorithms and produce research papers, yet understanding whether they exhibit creativity, the capacity to produce solutions that are both novel and useful, remains an open question. We present a framework for evaluating multi-turn LLM research agents' creativity using ML engineering tasks as a testbed, through three dimensions: P-Creativity (psychological novelty: novel relative to the agent's own prior solutions within a run), H-Creativity (historical novelty: novel relative to the corpus of human solutions), and Usefulness (task performance). Evaluating two agent frameworks, AIDE and AIRA-Dojo, on 10 Kaggle-style machine learning tasks from MLE-Bench, we develop an LLM-as-a-Judge pipeline and verify its strong correlation with human creativity judgments, providing a reliable automated metric for P-Creativity evaluation at scale. Applying this pipeline to agent trajectories, we find: (1) all agents exhibit declining P-Creativity as they transition from exploration to exploitation; (2) LLMs exhibit greater H-Creativity than medal-winning humans, yet achieve lower performance. Our findings reveal that current agents can explore novel regions of the solution space but lack the capacity to convert this novelty into improved task performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。