arXiv:2510.07307cs.LGcs.AI2025-10被引 6

用自动化多智能体流水线生成高质量可验证的机器学习任务

MLE-Smith: Scaling MLE Tasks with Automated Multi-Agent Pipeline

  • 构建生成-验证-执行的自动化流程,自动生成多样化机器学习任务
  • 在224个真实数据集上生成606个任务,覆盖多类别与多模态
  • 生成任务与人工设计任务表现高度相关,适合评估大模型工程能力

尽管语言模型在自动化机器学习工程方面取得进展,但高质量训练数据的获取仍受限。现有基准依赖静态人工任务,难以扩展且应用有限。本文提出MLE-Smith,一个全自动多智能体流水线,通过生成-验证-执行范式,将原始数据转化为竞赛式机器学习挑战,实现任务规模扩展、质量可验证、真实可用且多样性丰富。该流水线推动结构化任务设计与标准化重构,并采用混合验证机制,确保结构严格性和语义合理性;通过交互式执行验证实际可解性与现实契合度。我们在224个真实数据集上生成606个任务,涵盖多种类别、目标与模态,验证了其跨数据集的有效性。评估显示,8个主流及前沿大模型在这些任务上的表现与人工设计任务高度相关,证明MLE-Smith能有效扩展机器学习任务并保持质量。

原文摘要 · Abstract (English)

While Language Models (LMs) have made significant progress in automating machine learning engineering (MLE), the acquisition of high-quality MLE training data is significantly constrained. Current MLE benchmarks suffer from low scalability and limited applicability because they rely on static, manually curated tasks, demanding extensive time and manual effort to produce. We introduce MLE-Smith, a fully automated multi-agent pipeline, to transform raw datasets into competition-style MLE challenges through an efficient generate-verify-execute paradigm for scaling MLE tasks with verifiable quality, real-world usability, and rich diversity. The proposed multi-agent pipeline in MLE-Smith drives structured task design and standardized refactoring, coupled with a hybrid verification mechanism that enforces strict structural rules and high-level semantic soundness. It further validates empirical solvability and real-world fidelity through interactive execution. We apply MLE-Smith to 224 of real-world datasets and generate 606 tasks spanning multiple categories, objectives, and modalities, demonstrating that MLE-Smith can work effectively across a wide range of real-world datasets. Evaluation on the generated tasks shows that the performance of eight mainstream and cutting-edge LLMs on MLE-Smith tasks is strongly correlated with their performance on carefully human-designed tasks, highlighting the effectiveness of the MLE-Smith to scaling up MLE tasks, while maintaining task quality.

机器学习自动化多智能体任务生成大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。