通过划分数据空间提升模型精度,关键在于量化数据异质性。
Divide and Predict: An Architecture for Input Space Partitioning and Enhanced Accuracy
- 用成对样本影响的方差衡量数据异质性,识别混合分布。
- 等比例混合时方差最大,且基于此净化数据可显著提纯准确率。
- 适合需要提升复杂数据上模型性能的研究者使用。
本文提出一种内在度量方法,用于量化监督学习中训练数据的异质性。该度量基于一个随机变量的方差,其影响因素来自训练样本对之间的相互作用。证明该方差能有效捕捉数据异质性,可用于判断样本是否为多个分布的混合体。作者进一步证实数据本身蕴含支持分块划分的关键信息。通过在EMNIST图像数据和合成数据上的若干概念验证实验,表明当各分布比例相等时方差达到最大值。并展示了基于方差进行数据净化后,在各块上进行常规训练,可显著提升测试准确率。
原文摘要 · Abstract (English)
In this article the authors develop an intrinsic measure for quantifying heterogeneity in training data for supervised learning. This measure is the variance of a random variable which factors through the influences of pairs of training points. The variance is shown to capture data heterogeneity and can thus be used to assess if a sample is a mixture of distributions. The authors prove that the data itself contains key information that supports a partitioning into blocks. Several proof of concept studies are provided that quantify the connection between variance and heterogeneity for EMNIST image data and synthetic data. The authors establish that variance is maximal for equal mixes of distributions, and detail how variance-based data purification followed by conventional training over blocks can lead to significant increases in test accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。