arXiv:2505.17465cs.CLstat.ME2025-05EMNLP被引 3

提出自动生成机器学习排行榜的统一框架,解决数据不一致难题。

A Position Paper on the Automatic Generation of Machine Learning Leaderboards

  • 构建统一概念框架,规范自动排行榜生成任务定义
  • 建议使用标准化数据集和评估指标,提升可复现性
  • 推动覆盖全部结果与丰富元数据,助力研究透明化

机器学习研究中,比较已有工作通常依赖排行榜——以表格形式呈现相同任务、数据集和指标下的实验结果。然而,文献数量激增使得人工维护排行榜日益困难。为减轻负担,已有研究尝试从论文中自动提取排行榜条目。但现有方法在问题定义、范围和输出格式上差异显著,难以比较且实用性受限。本文首次系统梳理自动排行榜生成(ALG)研究,识别其核心假设、范围与输出形式的差异,提出统一的概念框架以标准化任务定义。同时给出基准测试指南,推荐促进公平、可复现评估的数据集与指标。最后,展望未来挑战与方向,如主张纳入所有报告结果及更丰富的元数据信息。

原文摘要 · Abstract (English)

An important task in machine learning (ML) research is comparing prior work, which is often performed via ML leaderboards: a tabular overview of experiments with comparable conditions (e.g., same task, dataset, and metric). However, the growing volume of literature creates challenges in creating and maintaining these leaderboards. To ease this burden, researchers have developed methods to extract leaderboard entries from research papers for automated leaderboard curation. Yet, prior work varies in problem framing, complicating comparisons and limiting real-world applicability. In this position paper, we present the first overview of Automatic Leaderboard Generation (ALG) research, identifying fundamental differences in assumptions, scope, and output formats. We propose an ALG unified conceptual framework to standardise how the ALG task is defined. We offer ALG benchmarking guidelines, including recommendations for datasets and metrics that promote fair, reproducible evaluation. Lastly, we outline challenges and new directions for ALG, such as, advocating for broader coverage by including all reported results and richer metadata.

机器学习排行榜自动化评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。