用特征效应大小预测模型性能和收敛速度,发现效果不佳。
Exploring the Impact of Dataset Statistical Effect Size on Model Performance and Data Sample Size Sufficiency
- 用特征效应大小衡量类别差异,看能否预判模型表现
- 实验显示效应大小与收敛速度、所需样本量无明显关联
- 提示需探索新方法来提前评估数据是否足够
拥有足够数量的高质量数据是训练有效机器学习模型的关键。若能在训练前有效判断数据集是否充足,将极大助力实验设计与数据收集。然而,这一能力至今仍难以实现。本文通过两项实验,探讨基本描述性统计量是否可作为数据有效性指标。首先,考察特征效应大小与模型性能之间的相关性(假设类别间差异越大,分类器表现越好);其次,分析效应大小是否影响学习曲线的收敛速度(假设效应越大,模型收敛越快,所需样本更少)。结果表明,效应大小无法有效预测模型性能或样本量需求,说明当前方法不足以前瞻性评估数据充足性,仍需进一步研究。
原文摘要 · Abstract (English)
Having a sufficient quantity of quality data is a critical enabler of training effective machine learning models. Being able to effectively determine the adequacy of a dataset prior to training and evaluating a model's performance would be an essential tool for anyone engaged in experimental design or data collection. However, despite the need for it, the ability to prospectively assess data sufficiency remains an elusive capability. We report here on two experiments undertaken in an attempt to better ascertain whether or not basic descriptive statistical measures can be indicative of how effective a dataset will be at training a resulting model. Leveraging the effect size of our features, this work first explores whether or not a correlation exists between effect size, and resulting model performance (theorizing that the magnitude of the distinction between classes could correlate to a classifier's resulting success). We then explore whether or not the magnitude of the effect size will impact the rate of convergence of the learning curve, (theorizing again that a greater effect size may indicate that the model will converge more rapidly, and with a smaller sample size needed). Our results appear to indicate that this is not an effective heuristic for determining adequate sample size or projecting model performance, and therefore that additional work is still needed to better prospectively assess adequacy of data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。