arXiv:2603.24999stat.APcs.AI2026-03

用新方法快速识别评估题中的坏题目,无需假设模型。

Efficient Detection of Bad Benchmark Items with Novel Scalability Coefficients

  • 基于单调回归的非参数系数,捕捉题目间单调关系。
  • 在多个数据集上优于或持平传统方法,AUC表现顶尖。
  • 轻量高效,适合大规模评测,支持多种题型。

评估的有效性依赖于题目质量,但现代评测工具常包含数千道题目且缺乏心理测量检验。本文提出一类基于项目间保序回归的新非参数可扩展性系数,用于高效检测全局劣质题目(如错键、表述模糊或构念不符)。核心贡献是带符号的保序$R^2$,它衡量一个题目方差中可被另一个题目的单调函数解释的最大比例,并通过Kendall's $τ$保留关联方向。将这些成对系数聚合为题项目级得分,可在不假设线性或参数响应模型的前提下,清晰区分问题题目与合格题目。我们证明带符号的保序$R^2$在所有单调预测器中具有极值性(提取最强单调信号),且此最优性直接转化为实际筛查能力。在三个AI基准数据集(HS Math、GSM8K、MMLU)和两个人类评估数据集上,该方法始终达到顶级AUC表现,优于或匹配经典测试理论、项目反应理论及基于维度的诊断方法。关键优势在于:在小样本大变量条件下仍稳健,仅需秒级计算的二元单调拟合,且无需修改即可处理二元、有序、连续混合题型。这是一种轻量、无模型依赖的过滤器,可显著降低现代大规模评测中发现缺陷题目的审阅成本。

原文摘要 · Abstract (English)

The validity of assessments, from large-scale AI benchmarks to human classrooms, depends on the quality of individual items, yet modern evaluation instruments often contain thousands of items with minimal psychometric vetting. We introduce a new family of nonparametric scalability coefficients based on interitem isotonic regression for efficiently detecting globally bad items (e.g., miskeyed, ambiguously worded, or construct-misaligned). The central contribution is the signed isotonic $R^2$, which measures the maximal proportion of variance in one item explainable by a monotone function of another while preserving the direction of association via Kendall's $τ$. Aggregating these pairwise coefficients yields item-level scores that sharply separate problematic items from acceptable ones without assuming linearity or committing to a parametric item response model. We show that the signed isotonic $R^2$ is extremal among monotone predictors (it extracts the strongest possible monotone signal between any two items) and show that this optimality property translates directly into practical screening power. Across three AI benchmark datasets (HS Math, GSM8K, MMLU) and two human assessment datasets, the signed isotonic $R^2$ consistently achieves top-tier AUC for ranking bad items above good ones, outperforming or matching a comprehensive battery of classical test theory, item response theory, and dimensionality-based diagnostics. Crucially, the method remains robust under the small-n/large-p conditions typical of AI evaluation, requires only bivariate monotone fits computable in seconds, and handles mixed item types (binary, ordinal, continuous) without modification. It is a lightweight, model-agnostic filter that can materially reduce the reviewer effort needed to find flawed items in modern large-scale evaluation regimes.

评测优化题目筛选非参数方法AI评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。