arXiv:2505.22813cs.LG2025-05被引 4

数据集质量是内在属性,不受模型和样本影响。

X-Factor: Quality Is a Dataset-Intrinsic Property

  • 控制样本量与类别平衡,测试多种模型表现。
  • 不同模型在同数据集上表现相关性达0.79,证明质量独立于模型。
  • 数据集质量源于其类别的内在质量,适合优化数据筛选者。

在优化机器学习分类器的普遍努力中,模型架构、数据集规模和类别平衡三个因素已被证实影响测试性能,但无法完全解释。此前已有证据表明存在一个额外因素——数据集质量,但其究竟是数据集与模型的联合属性,还是数据集自身的固有属性尚不明确。若质量确实是数据集的内在属性,独立于模型架构、数据集规模和类别平衡,则同一数据集在不同条件下应始终表现更优(或更差)。为此,我们构建了数千个数据集,每个均控制规模与类别平衡,并使用从随机森林到深度网络的多种架构进行训练。结果发现,分类器性能在不同架构间具有强相关性(R²=0.79),支持质量是独立于规模、类别平衡及模型架构的数据集内在属性。进一步分析表明,数据集质量似乎是更基本成分——各构成类别的质量——的涌现特性。因此,质量与规模、类别平衡、模型架构并列,成为影响性能的独立因素,也是优化机器学习分类的重要目标。

原文摘要 · Abstract (English)

In the universal quest to optimize machine-learning classifiers, three factors -- model architecture, dataset size, and class balance -- have been shown to influence test-time performance but do not fully account for it. Previously, evidence was presented for an additional factor that can be referred to as dataset quality, but it was unclear whether this was actually a joint property of the dataset and the model architecture, or an intrinsic property of the dataset itself. If quality is truly dataset-intrinsic and independent of model architecture, dataset size, and class balance, then the same datasets should perform better (or worse) regardless of these other factors. To test this hypothesis, here we create thousands of datasets, each controlled for size and class balance, and use them to train classifiers with a wide range of architectures, from random forests and support-vector machines to deep networks. We find that classifier performance correlates strongly by subset across architectures ($R^2=0.79$), supporting quality as an intrinsic property of datasets independent of dataset size and class balance and of model architecture. Digging deeper, we find that dataset quality appears to be an emergent property of something more fundamental: the quality of datasets' constituent classes. Thus, quality joins size, class balance, and model architecture as an independent correlate of performance and a separate target for optimizing machine-learning-based classification.

数据质量分类器优化模型无关

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。