arXiv:2506.23419cs.LGcs.AI2025-06被引 3

将任意科学数据集转化为可复现基准的自动化工具

BenchMake: Turn any scientific data set into a reproducible benchmark

  • 用非负矩阵分解识别凸包边缘的难题样本
  • 按需划分测试集,最大化差异性和统计显著性
  • 支持表格、图像、文本等多模态数据,适合科研评估

基准数据集是机器学习发展与应用的基石,确保新方法具备鲁棒性、可靠性与竞争力。然而,由于计算科学问题的独特性及领域快速演进,基准集相对稀少,导致计算科学家难以评估新成果。本文提出并测试了一种新工具 BenchMake,旨在将日益增多的公开科学数据集转化为社区可用的可复现基准。BenchMake 利用非负矩阵分解,确定性地识别并分离凸包(包含所有现有数据实例的最小凸集)上的挑战性边缘案例,并将指定比例的匹配数据实例划分为测试集,以在表格、图、图像、信号和文本等多种模态中实现最大差异性和统计显著性。通过对比十组来自不同科学领域的公开基准集(涵盖不同规模、形状与分布),验证了 BenchMake 分割优于传统划分与随机分割。

原文摘要 · Abstract (English)

Benchmark data sets are a cornerstone of machine learning development and applications, ensuring new methods are robust, reliable and competitive. The relative rarity of benchmark sets in computational science, due to the uniqueness of the problems and the pace of change in the associated domains, makes evaluating new innovations difficult for computational scientists. In this paper a new tool is developed and tested to potentially turn any of the increasing numbers of scientific data sets made openly available into a benchmark accessible to the community. BenchMake uses non-negative matrix factorisation to deterministically identify and isolate challenging edge cases on the convex hull (the smallest convex set that contains all existing data instances) and partitions a required fraction of matched data instances into a testing set that maximises divergence and statistical significance, across tabular, graph, image, signal and textual modalities. BenchMake splits are compared to establish splits and random splits using ten publicly available benchmark sets from different areas of science, with different sizes, shapes, distributions.

数据集基准评估可复现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。