arXiv:2508.10628cs.LG2025-08中稿 · the ENIAC 2025 con…被引 1

用心理测量学方法提升数据划分质量,让模型评估更精准。

Beyond Random Sampling: Instance Quality-Based Data Partitioning via Item Response Theory

  • 基于项目反应理论构建实例质量评分,指导数据划分
  • 高猜测性样本训练导致准确率低于50%,其他划分超70%
  • 可发现数据集中隐藏的有意义子群体,适合模型调试

机器学习模型的稳健验证至关重要,但传统数据划分方法常忽略实例本身的内在质量。本文提出利用项目反应理论(IRT)参数来表征并指导模型验证阶段的数据划分。在四个表格数据集上评估了基于IRT的划分策略对多种机器学习模型性能的影响。结果表明,IRT揭示了实例间的固有异质性,并识别出同一数据集中具有信息量的子群体。基于IRT生成的平衡划分能持续帮助理解模型偏差与方差之间的权衡。此外,猜测参数起决定性作用:使用高猜测性实例训练会使模型性能显著下降,出现准确率低于50%的情况,而其他划分方式在同一数据集上可达70%以上。

原文摘要 · Abstract (English)

Robust validation of Machine Learning (ML) models is essential, but traditional data partitioning approaches often ignore the intrinsic quality of each instance. This study proposes the use of Item Response Theory (IRT) parameters to characterize and guide the partitioning of datasets in the model validation stage. The impact of IRT-informed partitioning strategies on the performance of several ML models in four tabular datasets was evaluated. The results obtained demonstrate that IRT reveals an inherent heterogeneity of the instances and highlights the existence of informative subgroups of instances within the same dataset. Based on IRT, balanced partitions were created that consistently help to better understand the tradeoff between bias and variance of the models. In addition, the guessing parameter proved to be a determining factor: training with high-guessing instances can significantly impair model performance and resulted in cases with accuracy below 50%, while other partitions reached more than 70% in the same dataset.

数据划分项目反应理论模型评估机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。