arXiv:2411.01280cs.CL2024-11被引 2

用NLP自动评估填空题答案语义匹配度,提升大规模阅读测试效率

NLP and Education: using semantic similarity to evaluate filled gaps in a large-scale Cloze test in the classroom

  • 通过词嵌入模型计算答案与标准答案的语义相似度
  • 巴西语GloVe模型与人工评分相关性最高,达0.72
  • 适合教育机构批量评估学生阅读理解能力

本研究探讨了填空题(Cloze test)在大规模教学评估中的适用性及其挑战。为解决人工批改效率低的问题,提出一种基于自然语言处理(NLP)的自动化评分方法,利用词嵌入(WE)模型计算学生答案与预期答案之间的语义相似度。实验采用巴西葡萄牙语(PT-BR)的WE模型,分析来自巴西学生的真实填空数据。通过12名评审员对答案进行人工分类,对比模型评分与人工判断结果。结果显示,GloVe模型在语义相似度评估中表现最佳,与人工评分的相关系数达0.72,显著高于其他模型。研究表明,词嵌入模型可有效辅助大规模填空题评估,为教育测评提供了更高效、客观的技术路径。

原文摘要 · Abstract (English)

This study examines the applicability of the Cloze test, a widely used tool for assessing text comprehension proficiency, while highlighting its challenges in large-scale implementation. To address these limitations, an automated correction approach was proposed, utilizing Natural Language Processing (NLP) techniques, particularly word embeddings (WE) models, to assess semantic similarity between expected and provided answers. Using data from Cloze tests administered to students in Brazil, WE models for Brazilian Portuguese (PT-BR) were employed to measure the semantic similarity of the responses. The results were validated through an experimental setup involving twelve judges who classified the students' answers. A comparative analysis between the WE models' scores and the judges' evaluations revealed that GloVe was the most effective model, demonstrating the highest correlation with the judges' assessments. This study underscores the utility of WE models in evaluating semantic similarity and their potential to enhance large-scale Cloze test assessments. Furthermore, it contributes to educational assessment methodologies by offering a more efficient approach to evaluating reading proficiency.

教育评估语义相似度NLP应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。