arXiv:2505.13291cs.LGcs.AI2025-05被引 15

构建可扩展的时间序列机器学习工程智能体评测框架

TimeSeriesGym: A Scalable Benchmark for (Time Series) Machine Learning Engineering Agents

  • 融合多领域任务,评估数据处理、代码翻译等综合能力
  • 支持提交文件、代码、模型等多种成果的多维评估
  • 兼顾量化指标与大模型判断,适配真实工程场景

我们提出TimeSeriesGym,一个可扩展的基准框架,用于评估人工智能(AI)智能体在时间序列机器学习工程挑战中的表现。现有基准缺乏可扩展性,仅聚焦于定义明确环境下的模型构建,且仅评估有限的研究成果(如CSV提交文件)。为使智能体评测更贴近机器学习工程实践,我们的框架在两个关键维度上实现扩展:首先,认识到有效机器学习工程需具备多样化技能,TimeSeriesGym整合来自多个领域和任务的挑战,涵盖数据处理、理解研究仓库、代码翻译等独立能力及其组合,并开发工具支持大规模设计多种挑战;其次,实现对多种研究成果(包括提交文件、代码、模型)的评估机制,结合精确数值指标与灵活的LLM评估方法。这一双重策略平衡了客观评价与上下文判断。尽管初始重点为时间序列应用,该框架可轻松扩展至其他数据模态,显著提升智能体评估的全面性与实用性。我们已开源该基准框架,以促进对AI智能体机器学习工程能力的后续研究。

原文摘要 · Abstract (English)

We introduce TimeSeriesGym, a scalable benchmarking framework for evaluating Artificial Intelligence (AI) agents on time series machine learning engineering challenges. Existing benchmarks lack scalability, focus narrowly on model building in well-defined settings, and evaluate only a limited set of research artifacts (e.g., CSV submission files). To make AI agent benchmarking more relevant to the practice of machine learning engineering, our framework scales along two critical dimensions. First, recognizing that effective ML engineering requires a range of diverse skills, TimeSeriesGym incorporates challenges from diverse sources spanning multiple domains and tasks. We design challenges to evaluate both isolated capabilities (including data handling, understanding research repositories, and code translation) and their combinations, and rather than addressing each challenge independently, we develop tools that support designing multiple challenges at scale. Second, we implement evaluation mechanisms for multiple research artifacts, including submission files, code, and models, using both precise numeric measures and more flexible LLM-based evaluation approaches. This dual strategy balances objective assessment with contextual judgment. Although our initial focus is on time series applications, our framework can be readily extended to other data modalities, broadly enhancing the comprehensiveness and practical utility of agentic AI evaluation. We open-source our benchmarking framework to facilitate future research on the ML engineering capabilities of AI agents.

时间序列智能体评测机器学习工程可扩展性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。