用文本缺陷分析预判题目难易度,提升测验题质量筛选效率
The Impact of Item-Writing Flaws on Difficulty and Discrimination in Item Response Theory
- 基于19项文本缺陷标准自动标注7126道理工科选择题
- 缺陷越多,题目越容易且区分度越低,尤其在生命/地球科学中明显
- 可快速筛查低质量题目,适合大规模题库初筛
高质量测试题对教育评估至关重要,尤其在项目反应理论(IRT)框架下。传统验证依赖耗时的试测来估计题目难度与区分度。近年来,基于文本特征的题目编写缺陷(IWF)评分体系作为通用方法出现,可在无学生数据前提下实现可扩展的部署前评估,但其对实际IRT参数的预测效度尚不明确。为此,本研究分析了来自物理科学、数学及生命/地球科学领域的7,126道多项选择题,采用自动化方法对其标注19项IWF标准,并探究其与数据驱动的IRT参数间的关系。结果表明,IWF数量与题目难度及区分度存在显著统计关联,尤其在生命/地球科学与物理科学领域更为突出。研究还发现不同缺陷类型影响程度各异(如否定表述比不合理干扰项影响更严重),并揭示其如何使题目变得更难或更易。总体而言,自动化IWF分析可作为传统验证的有效补充,为题库初期筛选提供高效手段,尤其适用于识别低难度题目。研究呼吁进一步探索通用评估体系与理解领域知识的算法,以实现更稳健的题目验证。
原文摘要 · Abstract (English)
High-quality test items are essential for educational assessments, particularly within Item Response Theory (IRT). Traditional validation methods rely on resource-intensive pilot testing to estimate item difficulty and discrimination. More recently, Item-Writing Flaw (IWF) rubrics emerged as a domain-general approach for evaluating test items based on textual features. This method offers a scalable, pre-deployment evaluation without requiring student data, but its predictive validity concerning empirical IRT parameters is underexplored. To address this gap, we conducted a study involving 7,126 multiple-choice questions across various STEM subjects (physical science, mathematics, and life/earth sciences). Using an automated approach, we annotated each question with a 19-criteria IWF rubric and studied relationships to data-driven IRT parameters. Our analysis revealed statistically significant links between the number of IWFs and IRT difficulty and discrimination parameters, particularly in life/earth and physical science domains. We further observed how specific IWF criteria can impact item quality more and less severely (e.g., negative wording vs. implausible distractors) and how they might make a question more or less challenging. Overall, our findings establish automated IWF analysis as a valuable supplement to traditional validation, providing an efficient method for initial item screening, particularly for flagging low-difficulty MCQs. Our findings show the need for further research on domain-general evaluation rubrics and algorithms that understand domain-specific content for robust item validation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。