新基准FML-bench评估机器学习科研代理的探索能力与研究效果。
FML-bench: Benchmarking Machine Learning Agents for Scientific Research
- 设计8项基础科研任务,聚焦探索过程而非仅结果
- 提出探索多样性指标,揭示探索模式影响研究成效
- 验证广泛探索策略能提升性能,适合科研自动化研究者
大语言模型激发了对可自主提出想法并开展实验的机器学习科研代理的研究兴趣。然而,现有基准多从工程角度出发,侧重应用任务和最终性能、计算成本,忽视了代理在科研过程中的表现,限制了对其科研能力的全面评估。为此,我们提出FML-bench,包含8个多样且基础的机器学习研究任务,并引入互补性评估指标,特别是探索多样性(Exploration Diversity),用于量化迭代中提案的差异程度,揭示探索模式如何影响研究结果。我们在FML-bench上评估了先进科研代理,发现采用广泛探索策略的代理具有更高探索多样性,并在多个任务中取得更优性能;探索多样性与性能提升呈正相关。我们希望这些发现及基准能推动未来代理设计,并支持社区深入研究代理行为。基准代码已开源:https://github.com/qrzou/FML-bench。
原文摘要 · Abstract (English)
Large language models (LLMs) have sparked growing interest in machine learning research agents that can autonomously propose ideas and conduct experiments. However, existing benchmarks predominantly adopt an engineering-oriented perspective: they emphasize application-oriented tasks and evaluate primarily on final performance and computational cost, overlooking agents' research processes and limiting assessment of their capabilities in scientific research settings. To more comprehensively evaluate agents in scientific research settings, we introduce FML-bench, a benchmark comprising 8 diverse and fundamental ML research tasks, and further propose complementary metrics, notably Exploration Diversity, which quantifies the variance of proposals across iterations and reveals how exploration patterns influence research outcomes. We evaluate state-of-the-art research agents on FML-bench, showing that agents employing broad exploration strategies exhibit higher exploration diversity and achieve superior performance, and that exploration diversity positively correlates with performance improvements across multiple tasks. We hope these findings and our benchmark inform future agent design and support the community in further investigating agent behavior. Our benchmark is available at https://github.com/qrzou/FML-bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。