用心理测量学框架提升视觉模型与数据集的评估深度
On Evaluation of Vision Datasets and Models using Human Competency Frameworks
- 引入项目反应理论,为模型和数据项推断可解释的潜在参数
- 实现模型校准分析与高价值数据子集筛选,超越单一准确率
- 适合关注模型评估方法论、数据质量分析的研究者
计算机视觉中的模型与数据集评估仍具挑战性,现有排行榜多依赖单一准确率。尽管准确率是常用指标,但仅提供粗略评估,忽略了模型在所有数据项上的表现差异。本文探索项目反应理论(Item Response Theory, IRT),该框架可为一组模型及每个数据项推断出可解释的潜在参数,从而实现更丰富的评估与分析,超越单一准确率的局限。基于IRT,我们实现了模型校准评估、关键数据子集选择,并验证了其潜在参数在视觉模型与数据集比较分析中的有效性。
原文摘要 · Abstract (English)
Evaluating models and datasets in computer vision remains a challenging task, with most leaderboards relying solely on accuracy. While accuracy is a popular metric for model evaluation, it provides only a coarse assessment by considering a single model's score on all dataset items. This paper explores Item Response Theory (IRT), a framework that infers interpretable latent parameters for an ensemble of models and each dataset item, enabling richer evaluation and analysis beyond the single accuracy number. Leveraging IRT, we assess model calibration, select informative data subsets, and demonstrate the usefulness of its latent parameters for analyzing and comparing models and datasets in computer vision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。