研究训练数据质量对分类器性能的影响,发现数据退化会导致模型集体误判。
Effects of Training Data Quality on Classifier Performance
- 通过多种退化方式测试四类分类器在宏基因组数据上的表现
- 数据质量下降时,所有分类器从基本正确转为偶然正确,且错误趋同
- 数据与分析样本差异越大,分类结果越差,模型间一致性反而上升
我们通过大量数值实验评估并量化了分类器性能与训练数据质量的关系,这一因素常被忽视。在短DNA序列拼接成“contigs”的科学背景下,我们考察了多种机制导致的训练数据质量下降对四种分类器(贝叶斯分类器、神经网络、分段模型、随机森林)的影响,并分析了它们的个体行为与一致性。结果发现,随着数据退化,所有分类器均呈现类似“崩溃”现象:从多数正确变为仅偶然正确,且错误模式趋于一致。过程中显现出空间异质性特征:当训练数据远离分析数据时,分类决策质量下降,边界密度降低,但分类器间一致性增强。
原文摘要 · Abstract (English)
We describe extensive numerical experiments assessing and quantifying how classifier performance depends on the quality of the training data, a frequently neglected component of the analysis of classifiers. More specifically, in the scientific context of metagenomic assembly of short DNA reads into "contigs," we examine the effects of degrading the quality of the training data by multiple mechanisms, and for four classifiers -- Bayes classifiers, neural nets, partition models and random forests. We investigate both individual behavior and congruence among the classifiers. We find breakdown-like behavior that holds for all four classifiers, as degradation increases and they move from being mostly correct to only coincidentally correct, because they are wrong in the same way. In the process, a picture of spatial heterogeneity emerges: as the training data move farther from analysis data, classifier decisions degenerate, the boundary becomes less dense, and congruence increases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。