首个跨语言机器学习流程生成评测基准,测试大模型在多语种任务下的表现。
ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation
- 构建35个竞赛的14种语言版本,共490个任务-语言组合。
- 多语言性能下降因任务而异,部分场景出现严重退化。
- 适合评估大模型跨语言能力的研究者和开发者使用。
我们提出ML2B,首个针对大语言模型在跨语言任务理解与端到端机器学习流程生成方面进行系统评估的基准。尽管全球人工智能应用日益普及,但针对英语之外的任务描述,尚无系统性评估方法。ML2B通过将35个Kaggle竞赛(涵盖表格、文本、图像领域)由具备机器学习背景的母语研究人员翻译成14种语言,构建了490个任务-语言对。为保障评估可靠性,基准包含10个无公开解的私有竞赛,并采用网络隔离的评估环境,仅允许访问必要的机器学习资源。我们提供标准化评估协议、AutoGluon算法基线及全面的失败模式分析。对前沿模型(GPT-4.1-mini、GPT-OSS-120b、Gemini-2.5-Flash)的实验表明,跨语言性能下降高度依赖任务特性,而非遵循传统的资源可用性层级,性能差距从语言优势到严重退化不等。这一发现挑战了关于多语言模型能力的传统假设,强调了对机器学习流程生成进行系统性跨语言评估的必要性。相关代码、基线和评估框架已开源至https://github.com/enaix/ml2b。
原文摘要 · Abstract (English)
We introduce ML2B, the first benchmark for evaluating cross-lingual task comprehension in end-to-end ML pipeline generation by large language models. Despite growing global AI adoption, no systematic evaluation exists for ML pipeline generation beyond English task descriptions. ML2B addresses this gap with 35 Kaggle competitions spanning tabular, text, and image domains, translated into 14 languages by native-speaker researchers with ML expertise, yielding 490 task-language pairs. To ensure evaluation integrity, the benchmark incorporates 10 private competitions without publicly available solutions and employs network-isolated evaluation infrastructure restricting runtime access to essential ML resources. We provide standardized evaluation protocols, an AutoGluon algorithmic baseline, and comprehensive failure mode analysis. Experiments with frontier models (GPT-4.1-mini, GPT-OSS-120b, Gemini-2.5-Flash) reveal that cross-lingual performance degradation is highly task-dependent rather than following traditional resource-availability hierarchies, with gaps ranging from language advantages to severe degradation depending on competition characteristics. These findings challenge conventional assumptions about multilingual model capabilities and underscore the necessity of systematic cross-lingual evaluation for ML pipeline generation. We open-source the benchmark, baselines, and evaluation infrastructure at https://github.com/enaix/ml2b.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。