用大模型自动建科学排行榜,解决信息不全和错误问题。
Efficient Performance Tracking: Leveraging Large Language Models for Automated Construction of Scientific Leaderboards
- 用人工校验数据构建新榜单数据集SciLead,纠正旧数据缺陷。
- 在三种真实场景下测试,大模型能准找任务-数据-指标三元组。
- 适合做自动化评测系统或科研分析的研究者参考。
科学排行榜是标准化的排名体系,用于评估和比较竞争性方法。通常由任务、数据集和评估指标(TDM)三元组定义,支持客观性能评估并推动基准测试创新。然而,论文数量激增使手动构建和维护排行榜变得不可行。自动排行榜构建成为降低人工成本的解决方案。现有数据集基于社区贡献的排行榜,缺乏额外校验。我们分析发现大量排行榜信息不全,部分含错误。本文提出SciLead——一个经过人工校验的科学排行榜数据集,解决了上述问题。基于该数据集,我们设计三种实验设置,模拟构建排行榜时TDM三元组完全定义、部分定义或未定义的真实场景。以往研究仅覆盖第一种,后两种更贴近实际应用。为应对这些差异,我们开发了全面的基于大语言模型的排行榜构建框架。实验表明,不同LLM在识别TDM三元组方面表现良好,但在从论文中提取结果数值时存在困难。代码与数据已公开。
原文摘要 · Abstract (English)
Scientific leaderboards are standardized ranking systems that facilitate evaluating and comparing competitive methods. Typically, a leaderboard is defined by a task, dataset, and evaluation metric (TDM) triple, allowing objective performance assessment and fostering innovation through benchmarking. However, the exponential increase in publications has made it infeasible to construct and maintain these leaderboards manually. Automatic leaderboard construction has emerged as a solution to reduce manual labor. Existing datasets for this task are based on the community-contributed leaderboards without additional curation. Our analysis shows that a large portion of these leaderboards are incomplete, and some of them contain incorrect information. In this work, we present SciLead, a manually-curated Scientific Leaderboard dataset that overcomes the aforementioned problems. Building on this dataset, we propose three experimental settings that simulate real-world scenarios where TDM triples are fully defined, partially defined, or undefined during leaderboard construction. While previous research has only explored the first setting, the latter two are more representative of real-world applications. To address these diverse settings, we develop a comprehensive LLM-based framework for constructing leaderboards. Our experiments and analysis reveal that various LLMs often correctly identify TDM triples while struggling to extract result values from publications. We make our code and data publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。