研究发现,多样化的创意生成是AI科研代理成功的关键。
What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
- 通过分析MLE-bench上的代理轨迹,发现创意多样性影响性能。
- 控制实验显示,创意多样性越高,代理表现越强。
- 结果在多种评估指标下均成立,适用于优化科研代理设计。
AI研究代理有望通过自动化机器学习模型的设计、实现和训练来加速科学进步。然而该领域仍处于初期阶段,影响代理轨迹成败的关键因素尚不明确。本文研究了创意多样性在代理性能中的作用。首先,我们在MLE-bench这一知名基准上分析了不同模型与代理框架的代理轨迹,发现不同模型和框架产生的创意多样性各异,且表现更优的代理具有更高的创意多样性。随后,我们开展受控实验,调节创意多样性程度,证明更高的多样性带来更强性能。最后,通过考察除标准奖牌评分外的其他评估指标,进一步验证了上述结论在多种评价体系下依然成立。
原文摘要 · Abstract (English)
AI research agents offer the promise to accelerate scientific progress by automating the design, implementation, and training of machine learning models. However, the field is still in its infancy, and the key factors driving the success or failure of agent trajectories are not fully understood. We examine the role that ideation diversity plays in agent performance. First, we analyse agent trajectories on MLE-bench, a well-known benchmark to evaluate AI research agents, across different models and agent scaffolds. Our analysis reveals that different models and agent scaffolds yield varying degrees of ideation diversity, and that higher-performing agents tend to have increased ideation diversity. Further, we run a controlled experiment where we modify the degree of ideation diversity, demonstrating that higher ideation diversity results in stronger performance. Finally, we strengthen our results by examining additional evaluation metrics beyond the standard medal-based scoring of MLE-bench, showing that our findings still hold across other agent performance metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。