arXiv:2602.10441cs.LGcs.AI2026-02被引 2

首个面向数据湖多表学习的基准测试,助力评估真实场景下的机器学习性能。

LakeMLB: Data Lake Machine Learning Benchmark

  • 构建针对数据湖中表合并与关联场景的基准测试框架。
  • 涵盖6个真实数据集,支持预训练、数据增强等3种学习范式。
  • 适合研究数据湖机器学习的学者与工业界开发者使用。

数据湖已成为大规模机器学习的基础平台,支持异构数据的灵活管理。尽管其重要性日益凸显,针对数据湖环境下机器学习性能的标准评估基准仍十分稀缺。为此,我们提出LakeMLB(数据湖机器学习基准),首个专为数据湖中多表学习设计的基准。该基准聚焦于合并(Union)与关联(Join)两类典型场景,提供覆盖多个领域的6个真实数据集,支持预训练、数据增强和特征增强三种代表性多表学习范式,并配备标准化的数据划分与评估协议。我们对前沿表格学习方法进行了广泛实验,揭示了不同数据湖场景下的性能表现差异。相关数据集与代码已开源,可访问https://github.com/zhengwang100/LakeMLB,推动数据湖生态中机器学习研究的严谨发展。

原文摘要 · Abstract (English)

Data lakes have become a fundamental platform for large-scale machine learning by enabling flexible management of heterogeneous data. Despite their growing importance, standardized benchmarks for evaluating machine learning performance in data lake environments remain scarce. To address this gap, we present LakeMLB (Data Lake Machine Learning Benchmark), the first benchmark designed for multi-table machine learning in data lakes. LakeMLB focuses on two representative scenarios, Union and Join, and provides six real-world datasets spanning diverse domains. It supports three representative multi-table learning paradigms: pre-training, data augmentation, and feature augmentation, together with standardized data splits and evaluation protocols. We conduct extensive experiments with state-of-the-art tabular learning methods and provide insights into their performance across different data lake scenarios. We release both datasets and code to facilitate rigorous research on machine learning in data lake ecosystems; the benchmark is available at https://github.com/zhengwang100/LakeMLB.

数据湖多表学习基准测试表格数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。