AI可有效评估英语阅读题内容效度,表现接近人类专家。
The Use of Artificial Intelligence Tools in Assessing Content Validity: A Comparative Study with Human Experts
- 用AI与人类专家对比评估25道B1级英语题的内容效度。
- AI与人类评分无显著差异,关键指标一致。
- 适合需高效评估的教育测试开发场景。
本研究考察了AI评估者在评估B1级英语阅读理解试题内容效度方面是否与人类专家表现一致。开发了一套包含25道多项选择题的测试,并由四位人类专家和四位AI评估者进行评分。结果显示,人类与AI评估者的评分无统计学显著差异,评价趋势相似。通过威尔科克斯森符号秩检验分析内容效度比(CVR)和项目内容效度指数(I-CVI),亦未发现显著差异。研究发现,在某些情况下,AI评估者可替代人类专家。但个别题目差异可能源于对评估标准的不同理解。提升语言清晰度并明确定义评估标准,有助于提高评估一致性。因此建议发展人机协同的混合评估系统。
原文摘要 · Abstract (English)
In this study, it was investigated whether AI evaluators assess the content validity of B1-level English reading comprehension test items in a manner similar to human evaluators. A 25-item multiple-choice test was developed, and these test items were evaluated by four human and four AI evaluators. No statistically significant difference was found between the scores given by human and AI evaluators, with similar evaluation trends observed. The Content Validity Ratio (CVR) and the Item Content Validity Index (I-CVI) were calculated and analyzed using the Wilcoxon Signed-Rank Test, with no statistically significant difference. The findings revealed that in some cases, AI evaluators could replace human evaluators. However, differences in specific items were thought to arise from varying interpretations of the evaluation criteria. Ensuring linguistic clarity and clearly defining criteria could contribute to more consistent evaluations. In this regard, the development of hybrid evaluation systems, in which AI technologies are used alongside human experts, is recommended.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。