用大模型自动匹配数学题与课程标准,准确率超95%。
Automated Alignment of Math Items to Content Standards in Large-Scale Assessments Using Language Models
- 用BERT类模型提取题目特征,自动分类到课程领域和技能点。
- DeBERTa-v3-base在领域分类上达95%准确率,RoBERTa-large在技能分类上达86.9%。
- 集成学习未超越单个最佳模型,线性模型效果受降维影响。
在大规模测评中,题目与课程标准的精准对齐对于分数解释的有效性至关重要。本研究评估了三种自动化方法,将题目映射到四个领域和十九个技能标签。首先,提取嵌入向量并训练多种经典监督学习模型,进一步探究降维对性能的影响;其次,微调八种BERT及其变体模型用于领域和技能对齐;第三,探索基于多数投票的集成学习和多模型元学习的堆叠方法。DeBERTa-v3-base在领域对齐上取得最高加权平均F1分数0.950,RoBERTa-large在技能对齐上达到最高F1分数0.869。集成模型未能超越表现最佳的语言模型。降维提升了基于嵌入的线性分类器性能,但整体仍不及语言模型。本研究展示了自动化题目对齐至课程标准的不同方法。
原文摘要 · Abstract (English)
Accurate alignment of items to content standards is critical for valid score interpretation in large-scale assessments. This study evaluates three automated paradigms for aligning items with four domain and nineteen skill labels. First, we extracted embeddings and trained multiple classical supervised machine learning models, and further investigated the impact of dimensionality reduction on model performance. Second, we fine-tuned eight BERT model and its variants for both domain and skill alignment. Third, we explored ensemble learning with majority voting and stacking with multiple meta-models. The DeBERTa-v3-base achieved the highest weighted-average F1 score of 0.950 for domain alignment while the RoBERTa-large yielded the highest F1 score of 0.869 for skill alignment. Ensemble models did not surpass the best-performing language models. Dimension reduction enhanced linear classifiers based on embeddings but did not perform better than language models. This study demonstrated different methods in automated item alignment to content standards.}
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。