arXiv:2509.09321cs.AI2025-09被引 5

用网页代理自动构建真实机器学习任务集,评估大模型智能体能力。

Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization

  • 通过浏览器自动化抓取Kaggle等平台任务,覆盖多模态数据
  • 基于参赛人数和分数分布建模难度,实现任务客观分级
  • 支持性能、格式、约束、泛化四维度评估,适合不同场景测试

大型语言模型(LLMs)的发展催生了可自动化完成端到端机器学习流程的通用智能体,涵盖数据分析、特征工程、模型训练与竞赛求解。然而现有基准在任务覆盖、领域多样性、难度建模和评估严谨性方面仍显不足,难以全面反映智能体在真实场景中的能力。本文提出TAM Bench,一个多样、真实且结构化的基准,用于评估基于LLM的智能体在端到端机器学习任务上的表现。其三大创新包括:(1)基于浏览器自动化与LLM的任务获取系统,自动从Kaggle、AIcrowd、Biendata等平台收集并结构化机器学习挑战,覆盖多种任务类型与数据模态(如表格、文本、图像、图、音频);(2)基于排行榜的难度建模机制,利用参与者数量与得分分散度估算任务复杂度,实现可扩展且客观的任务校准;(3)多维度评估框架,包含性能、格式合规性、约束遵守与任务泛化能力。基于150个精心筛选的AutoML任务,构建了三个不同规模的子集——Lite、Medium、Full,适用于不同评估场景。其中Lite版本含18个任务,各模态与难度水平均衡,适合作为日常评测与对比研究的实用基准。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) have enabled the emergence of general-purpose agents for automating end-to-end machine learning (ML) workflows, including data analysis, feature engineering, model training, and competition solving. However, existing benchmarks remain limited in task coverage, domain diversity, difficulty modeling, and evaluation rigor, failing to capture the full capabilities of such agents in realistic settings. We present TAM Bench, a diverse, realistic, and structured benchmark for evaluating LLM-based agents on end-to-end ML tasks. TAM Bench features three key innovations: (1) A browser automation and LLM-based task acquisition system that automatically collects and structures ML challenges from platforms such as Kaggle, AIcrowd, and Biendata, spanning multiple task types and data modalities (e.g., tabular, text, image, graph, audio); (2) A leaderboard-driven difficulty modeling mechanism that estimates task complexity using participant counts and score dispersion, enabling scalable and objective task calibration; (3) A multi-dimensional evaluation framework incorporating performance, format compliance, constraint adherence, and task generalization. Based on 150 curated AutoML tasks, we construct three benchmark subsets of different sizes -- Lite, Medium, and Full -- designed for varying evaluation scenarios. The Lite version, with 18 tasks and balanced coverage across modalities and difficulty levels, serves as a practical testbed for daily benchmarking and comparative studies.

智能体自动评估基准测试端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。