用大模型自动检查试题与课标匹配度,效率提升且准确率超九成。
Scaling Item-to-Standard Alignment with Large Language Models: Accuracy, Limits, and Solutions
- 大模型结合候选筛选策略,自动识别试题与课标是否匹配。
- 在数学题中准确率达94%,阅读题因标准重叠略低。
- 预筛候选技能后,正确答案前五名覆盖率达95%以上。
随着教育体系发展,确保测评题目与课程标准一致对公平性和教学相关性至关重要。传统人工审核虽准确但耗时费力,尤其面对大规模题库时。本研究检验大型语言模型(LLMs)能否在不牺牲准确性的前提下加速该过程。基于12,000+个K-5年级的题目-技能配对,测试了GPT-3.5 Turbo、GPT-4o-mini和GPT-4o在三项真实任务中的表现:识别错位题目、从全部标准中选出正确技能、分类前缩小候选列表。研究1显示,GPT-4o-mini在约83%-94%情况下正确判断对齐状态,包括细微偏差;研究2表明数学题表现良好,阅读题因标准语义重叠导致准确率下降;研究3证明预筛选候选技能可显著提升效果,正确技能在前五个建议中出现超过95%。结果表明,结合候选过滤策略的大模型可大幅减少人工审核负担,同时保持对齐准确性。建议构建混合流程,由大模型初步筛查,人工处理模糊案例,实现持续题目验证与教学对齐的规模化解决方案。
原文摘要 · Abstract (English)
As educational systems evolve, ensuring that assessment items remain aligned with content standards is essential for maintaining fairness and instructional relevance. Traditional human alignment reviews are accurate but slow and labor-intensive, especially across large item banks. This study examines whether Large Language Models (LLMs) can accelerate this process without sacrificing accuracy. Using over 12,000 item-skill pairs in grades K-5, we tested three LLMs (GPT-3.5 Turbo, GPT-4o-mini, and GPT-4o) across three tasks that mirror real-world challenges: identifying misaligned items, selecting the correct skill from the full set of standards, and narrowing candidate lists prior to classification. In Study 1, GPT-4o-mini correctly identified alignment status in approximately 83-94% of cases, including subtle misalignments. In Study 2, performance remained strong in mathematics but was lower for reading, where standards are more semantically overlapping. Study 3 demonstrated that pre-filtering candidate skills substantially improved results, with the correct skill appearing among the top five suggestions more than 95% of the time. These findings suggest that LLMs, particularly when paired with candidate filtering strategies, can significantly reduce the manual burden of item review while preserving alignment accuracy. We recommend the development of hybrid pipelines that combine LLM-based screening with human review in ambiguous cases, offering a scalable solution for ongoing item validation and instructional alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。