整理46个表格型序数分类数据集,助力方法可复现验证
TOC-UCO: a comprehensive repository of tabular ordinal classification datasets
- 构建统一预处理的46个表格型序数分类数据集
- 提供30组随机训练测试划分,支持实验可复现
- 适合开发新序数分类算法的研究者使用
序数分类(OC)是一种类别间具有自然顺序关系的特殊分类问题,在众多实际应用中存在。尽管近年来涌现大量序数分类方法,但该领域发展受限于缺乏统一的数据集基准。为此,科尔多瓦大学(UCO)推出TOC-UCO(UCO表格型序数分类数据集仓库),公开提供46个经过统一预处理的表格型序数分类数据集,每个数据集均具备足够样本量与合理类别分布。所有数据集来源与预处理流程均详细说明,并提供30组随机划分的训练测试集,便于新方法的标准化评估与结果复现。
原文摘要 · Abstract (English)
An ordinal classification (OC) problem corresponds to a special type of classification characterised by the presence of a natural order relationship among the classes. This type of problem can be found in a number of real-world applications, motivating the design and development of many ordinal methodologies over the last years. However, it is important to highlight that the development of the OC field suffers from one main disadvantage: the lack of a comprehensive set of datasets on which novel approaches to the literature have to be benchmarked. In order to approach this objective, this manuscript from the University of Córdoba (UCO), which have previous experience on the OC field, provides the literature with a publicly available repository of tabular data for a robust validation of novel OC approaches, namely TOC-UCO (Tabular Ordinal Classification repository of the UCO). Specifically, this repository includes a set of $46$ tabular ordinal datasets, preprocessed under a common framework and ensured to have a reasonable number of patterns and an appropriate class distribution. We also provide the sources and preprocessing steps of each dataset, along with details on how to benchmark a novel approach using the TOC-UCO repository. For this, indices for $30$ different randomised train-test partitions are provided to facilitate the reproducibility of the experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。