用大模型和可解释AI找出影响测验公平性的关键词。
Finding Words Associated with DIF: Predicting Differential Item Functioning using LLMs and Explainable AI
- 用大模型分析题目文本,预测项目功能差异(DIF)
- 发现关联DIF的词多为设计内含的小领域术语,非无关内容
- 适合用于题目编写时筛查风险词,提升测评公平性
我们微调并比较了多种基于编码器的Transformer大语言模型(LLM),以从题目文本中预测项目功能差异(DIF)。随后应用可解释人工智能(XAI)方法识别与DIF相关的具体词汇。数据包含42,180道面向3至11年级学生、用于英语语言艺术和数学终结性州级评估的题目。在八组焦点群体与参照群体的对比中,预测的R²值范围为0.04至0.32。研究发现,许多与DIF相关的词汇反映的是测试蓝图中本就设计好的细微子领域,而非应被剔除的与测验构念无关的内容,这可能解释了为何以往对DIF题目的定性审查常得不出明确结论。该方法可用于题目撰写过程中即时筛查与DIF相关的词汇进行修改,或通过突出文本中的关键词汇来辅助传统DIF分析结果的解读。未来扩展可提升资源有限的评估项目公平性,尤其适用于样本量不足的小型子群体。
原文摘要 · Abstract (English)
We fine-tuned and compared several encoder-based Transformer large language models (LLM) to predict differential item functioning (DIF) from the item text. We then applied explainable artificial intelligence (XAI) methods to these models to identify specific words associated with DIF. The data included 42,180 items designed for English language arts and mathematics summative state assessments among students in grades 3 to 11. Prediction $R^2$ ranged from .04 to .32 among eight focal and reference group pairs. Our findings suggest that many words associated with DIF reflect minor sub-domains included in the test blueprint by design, rather than construct-irrelevant item content that should be removed from assessments. This may explain why qualitative reviews of DIF items often yield confusing or inconclusive results. Our approach can be used to screen words associated with DIF during the item-writing process for immediate revision, or help review traditional DIF analysis results by highlighting key words in the text. Extensions of this research can enhance the fairness of assessment programs, especially those that lack resources to build high-quality items, and among smaller subpopulations where we do not have sufficient sample sizes for traditional DIF analyses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。