arXiv:2504.14232cs.AIcs.CL2025-04被引 24

用认知框架提升AI出题质量,让机器更懂教育目标

Assessing AI-Generated Questions' Alignment with Cognitive Frameworks in Educational Assessment

  • 将布卢姆分类法嵌入AI出题插件,按认知层级生成题目
  • DistilBERT模型准确率达91%,显著提升高阶思维题识别
  • 适合教育技术开发者与智能测评系统研究者

本研究评估了将布卢姆分类法(Bloom's Taxonomy)整合进OneClickQuiz——一个用于Moodle平台的AI驱动多选题自动生成插件的效果。研究构建了一个包含3691道题目的数据集,按布卢姆层级进行标注,并采用多项式逻辑回归、朴素贝叶斯、线性支持向量分类(Linear SVC)及基于DistilBERT的Transformer模型进行分类性能比较。结果表明,较高认知层级的问题普遍具有更长的长度、更高的Flesch-Kincaid Grade Level(FKGL)和词汇密度(LD),反映出更高认知要求。多项式逻辑回归在‘知识’层级表现最佳,但在高阶层级准确率下降;合并高阶类别后整体性能提升。朴素贝叶斯与线性SVC对低阶任务有效,但难以处理高阶任务。而DistilBERT模型达到最高性能,总体验证准确率达91%,显著改善了低阶与高阶认知层级的分类效果。研究证实将布卢姆分类法融入AI评估工具的潜力,凸显先进模型如DistilBERT在教育内容生成中的优势。

原文摘要 · Abstract (English)

This study evaluates the integration of Bloom's Taxonomy into OneClickQuiz, an Artificial Intelligence (AI) driven plugin for automating Multiple-Choice Question (MCQ) generation in Moodle. Bloom's Taxonomy provides a structured framework for categorizing educational objectives into hierarchical cognitive levels. Our research investigates whether incorporating this taxonomy can improve the alignment of AI-generated questions with specific cognitive objectives. We developed a dataset of 3691 questions categorized according to Bloom's levels and employed various classification models-Multinomial Logistic Regression, Naive Bayes, Linear Support Vector Classification (SVC), and a Transformer-based model (DistilBERT)-to evaluate their effectiveness in categorizing questions. Our results indicate that higher Bloom's levels generally correlate with increased question length, Flesch-Kincaid Grade Level (FKGL), and Lexical Density (LD), reflecting the increased complexity of higher cognitive demands. Multinomial Logistic Regression showed varying accuracy across Bloom's levels, performing best for "Knowledge" and less accurately for higher-order levels. Merging higher-level categories improved accuracy for complex cognitive tasks. Naive Bayes and Linear SVC also demonstrated effective classification for lower levels but struggled with higher-order tasks. DistilBERT achieved the highest performance, significantly improving classification of both lower and higher-order cognitive levels, achieving an overall validation accuracy of 91%. This study highlights the potential of integrating Bloom's Taxonomy into AI-driven assessment tools and underscores the advantages of advanced models like DistilBERT for enhancing educational content generation.

AI出题布卢姆分类法认知层级教育评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。