arXiv:2601.22420cs.LGcs.AI2026-01Conference of the …

MetaLead收录全量实验结果,提升机器学习评估透明度。

MetaLead: A Comprehensive Human-Curated Leaderboard Dataset for Transparent Reporting of Machine Learning Experiments

  • 人工标注所有实验结果,非仅最佳数据
  • 区分基线、方法与变体,支持类型化对比
  • 明确分离训练/测试集,适合跨领域评估

排行榜在机器学习领域对基准测试和进展追踪至关重要。然而,传统排行榜创建需大量人工投入。近年来虽有自动化尝试,但现有数据集仅包含每篇论文的最佳结果,且元数据有限。我们提出MetaLead,一个完全人工标注的机器学习排行榜数据集,涵盖所有实验结果,提升结果透明度,并包含额外元数据,如实验类型(基线、所提方法、方法变体),支持按实验类型进行对比分析;同时显式分离训练集与测试集,便于跨领域评估。其丰富结构为机器学习研究提供了更透明、更细致的评估资源。

原文摘要 · Abstract (English)

Leaderboards are crucial in the machine learning (ML) domain for benchmarking and tracking progress. However, creating leaderboards traditionally demands significant manual effort. In recent years, efforts have been made to automate leaderboard generation, but existing datasets for this purpose are limited by capturing only the best results from each paper and limited metadata. We present MetaLead, a fully human-annotated ML Leaderboard dataset that captures all experimental results for result transparency and contains extra metadata, such as the result experimental type: baseline, proposed method, or variation of proposed method for experiment-type guided comparisons, and explicitly separates train and test dataset for cross-domain assessment. This enriched structure makes MetaLead a powerful resource for more transparent and nuanced evaluations across ML research.

机器学习评估基准数据透明排行榜

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。