用心理测量学方法提升模型评估精度,区分表现相近的模型
Standing on the shoulders of giants
- 引入项目反应理论(IRT)补充混淆矩阵,分析实例潜在特征
- 97%置信度下,IRT得分与66%经典指标贡献不同
- 适合需精细比较模型性能的研究者使用
尽管对机器学习发展至关重要,但源自混淆矩阵的经典评估指标(如精确率、F1值)存在局限:仅提供量化结果,未考虑数据复杂性或预测质量。为克服此问题,近期研究引入心理测量学指标如项目反应理论(IRT),可从实例潜在特征层面进行评估。本文探讨如何利用IRT概念丰富混淆矩阵,以在性能相近的模型中识别最优者。研究表明,IRT并非取代传统指标,而是提供新的评估维度,揭示模型在特定实例上的细微行为差异。实验发现,在97%置信水平下,IRT得分与所分析的66%经典指标具有不同贡献。
原文摘要 · Abstract (English)
Although fundamental to the advancement of Machine Learning, the classic evaluation metrics extracted from the confusion matrix, such as precision and F1, are limited. Such metrics only offer a quantitative view of the models' performance, without considering the complexity of the data or the quality of the hit. To overcome these limitations, recent research has introduced the use of psychometric metrics such as Item Response Theory (IRT), which allows an assessment at the level of latent characteristics of instances. This work investigates how IRT concepts can enrich a confusion matrix in order to identify which model is the most appropriate among options with similar performance. In the study carried out, IRT does not replace, but complements classical metrics by offering a new layer of evaluation and observation of the fine behavior of models in specific instances. It was also observed that there is 97% confidence that the score from the IRT has different contributions from 66% of the classical metrics analyzed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。