arXiv:2509.26431cs.CL2025-09被引 2

用小模型自动对齐试题与课程标准,准确率超传统方法。

Text-Based Approaches to Item Alignment to Content Standards in Large-Scale Reading & Writing Tests

  • 微调小语言模型,按领域和技能双维度自动对齐试题。
  • 增加试题文本数据显著提升性能,优于单纯增大样本量。
  • 适合教育评估、智能命题系统开发者参考。

将试题与课程标准对齐是测试开发中获取内容效度证据的关键步骤,传统依赖人工专家判断,存在主观性和耗时问题。本研究探讨了微调小语言模型(SLMs)在大规模标准化读写测试中实现自动化试题对齐的性能,数据来自美国大学入学考试(SAT)与预备考试(PSAT)。模型分别在4个内容领域和10项技能层级上进行训练与评估。在两个测试数据集上,多指标评估显示,增加试题文本数据可显著提升模型表现,超越仅扩大样本量带来的改进。作为对比,使用多语言-E5-large-instruct模型生成的嵌入向量训练监督学习模型,结果表明微调后的SLMs始终优于基于嵌入的机器学习模型,尤其在细粒度技能对齐任务中优势明显。通过余弦相似度、KL散度及嵌入投影等语义相似性分析发现,部分技能在语义上过于接近,解释了误分类现象。

原文摘要 · Abstract (English)

Aligning test items to content standards is a critical step in test development to collect validity evidence based on content. Item alignment has typically been conducted by human experts. This judgmental process can be subjective and time-consuming. This study investigated the performance of fine-tuned small language models (SLMs) for automated item alignment using data from a large-scale standardized reading and writing test for college admissions. Different SLMs were trained for alignment at both domain and skill levels respectively with 10 skills mapped to 4 content domains. The model performance was evaluated in multiple criteria on two testing datasets. The impact of types and sizes of the input data for training was investigated. Results showed that including more item text data led to substantially better model performance, surpassing the improvements induced by sample size increase alone. For comparison, supervised machine learning models were trained using the embeddings from the multilingual-E5-large-instruct model. The study results showed that fine-tuned SLMs consistently outperformed the embedding-based supervised machine learning models, particularly for the more fine-grained skill alignment. To better understand model misclassifications, multiple semantic similarity analysis including pairwise cosine similarity, Kullback-Leibler divergence of embedding distributions, and two-dimension projections of item embeddings were conducted. These analyses consistently showed that certain skills in SAT and PSAT were semantically too close, providing evidence for the observed misclassification.

试题对齐小模型教育评估自然语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。