arXiv:2605.15015cs.AIcs.CL2026-05被引 1

小模型可本地部署生成教育测评题,效果接近大模型但需人工把关。

Small, Private Language Models as Teammates for Educational Assessment Design

论文配图:Small, Private Language Models as Teammates for Educational Assessment Design
图 1 · 摘自论文原文
  • 用小语言模型本地生成符合布鲁姆分类法的测验题
  • 小模型在关键质量维度上表现接近大模型,但评估存在系统性偏差
  • 适合关注隐私与资源受限场景的教育技术研究者

生成式AI在教育设计中日益应用,如大型语言模型(LLMs)能生成符合教学框架(如布鲁姆分类法)的测评题。然而,现有方法常依赖主观或有限的评估方式,多聚焦专有模型,且很少系统考察生成、评估或实际教育环境中的部署约束。与此同时,小型语言模型(SLMs)作为本地化替代方案,更契合隐私与资源限制,但其在测评任务中的有效性仍待深入探索。为此,本研究系统比较了LLMs与SLMs在测评题设计中的表现,采用可复现、基于教育学原理的指标评估生成质量,并分析模型评分与专家评分的一致性与可靠性。结果表明,SLMs在关键教育质量维度上表现良好,支持本地、隐私敏感的部署;但模型评估相较专家评分存在系统性不一致与偏差。研究证实语言模型应作为评估流程中的有限助手,强调人机协同的必要性,并通过考察质量、可靠性和部署适配性的权衡,推动自动化教育题生成的发展。

原文摘要 · Abstract (English)

Generative AI increasingly supports educational design tasks, e.g., through Large Language Models (LLMs), demonstrating the capability to design assessment questions that are aligned with pedagogical frameworks (e.g., Bloom's taxonomy). However, they often rely on subjective or limited evaluation methods; focus primarily on proprietary models; or rarely systematically examine generation, evaluation, or deployment constraints in real educational settings. Meanwhile, Small Language Models (SLMs) have emerged as local alternatives that better address privacy and resource limitations; yet their effectiveness for assessment tasks remains underexplored. To address this gap, we systematically compare LLMs and SLMs for assessment question design; evaluate generation quality across Bloom's taxonomy levels using reproducible, pedagogically grounded metrics; and further assess model-based judging against expert-informed evaluation by analyzing reliability and agreement patterns. Results show that SLMs achieve competitive performance across key pedagogically motivated quality dimensions while enabling local, privacy-sensitive deployment. However, model-based evaluations also exhibit systematic inconsistencies and bias relative to expert ratings. These findings provide evidence to posit language models as bounded assistants in assessment workflows; underscore the necessity of Human-in-the-Loop; and advance the automated educational question generation field by examining quality, reliability, and deployment-aware trade-offs.

教育AI小模型测评生成人机协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。