arXiv:2506.20640cs.AIcs.LG2025-06被引 5

CoMind让机器学习代理像人一样协作,从社区中学习并超越多数人类选手。

CoMind: Towards Community-Driven Agents for Machine Learning Engineering

  • 设计多代理系统并行探索,结合社区知识提升解决方案质量。
  • 在75个历史竞赛中获得36%奖牌率,领先现有水平。
  • 在8个实时竞赛中平均超越92.6%人类选手,多次进入前1%。

大型语言模型代理在自动化机器学习工程方面展现出潜力。然而,现有代理通常孤立地解决特定问题,未能与更广泛的科研社区互动,而人类研究者常通过分享知识获得洞见。为弥合这一差距,我们提出MLE-Live,一个用于评估代理与模拟Kaggle科研社区交流及利用集体知识能力的实时评测框架。基于此框架,我们构建了CoMind,一种多代理系统,可系统性地获取外部知识。CoMind采用迭代并行探索机制,同时发展多个解决方案,在探索广度与实现深度间取得平衡。在我们MLE-Live框架内的75个过往Kaggle竞赛中,CoMind实现了36%的奖牌率,创下新纪录。关键的是,在8个正在进行的实时竞赛中,其平均表现超越92.6%的人类参赛者,三次登上官方排行榜前5%,一次进入前1%。

原文摘要 · Abstract (English)

Large language model (LLM) agents show promise in automating machine learning (ML) engineering. However, existing agents typically operate in isolation on a given research problem, without engaging with the broader research community, where human researchers often gain insights and contribute by sharing knowledge. To bridge this gap, we introduce MLE-Live, a live evaluation framework designed to assess an agent's ability to communicate with and leverage collective knowledge from a simulated Kaggle research community. Building on this framework, we propose CoMind, a multi-agent system designed to systematically leverage external knowledge. CoMind employs an iterative parallel exploration mechanism, developing multiple solutions simultaneously to balance exploratory breadth with implementation depth. On 75 past Kaggle competitions within our MLE-Live framework, CoMind achieves a 36% medal rate, establishing a new state of the art. Critically, when deployed in eight live, ongoing competitions, CoMind outperforms 92.6% of human competitors on average, placing in the top 5% on three official leaderboards and the top 1% on one.

机器学习多智能体自动化竞赛优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。