arXiv:2508.20217cs.CLcs.AI2025-08被引 1

用提示工程让小模型生成更精准的中小学测验题

Prompting Strategies for Language Model-Based Item Generation in K-12 Education: Bridging the Gap Between Small and Large Language Models

  • 用结构化提示提升小模型生成质量,尤其结合思维链与顺序设计
  • 小模型经优化后比大模型零样本输出更符合教学目标
  • 适合教育科技开发者和考试命题人员参考

本研究探索使用语言模型自动生成(AIG)选择题以降低中小学形态评估题目的开发成本与不一致性。采用两阶段方法:首先比较微调过的中等规模模型(Gemma, 2B)与未调优的大模型(GPT-3.5, 175B);其次评估七种结构化提示策略,包括零样本、少样本、思维链、角色设定、顺序生成及组合方式。生成题目通过自动化指标与专家评分在五个维度进行评估,并使用训练于专家标注样本的GPT-4.1模拟大规模人工评分。结果表明,结构化提示,特别是思维链与顺序设计结合的方式,显著提升了Gemma的表现。Gemma整体生成的题目比GPT-3.5零样本输出更具构念一致性与教学适切性,提示设计是中等规模模型性能的关键。研究表明,在数据有限条件下,结构化提示与高效微调可有效增强中等模型的AIG能力。研究强调结合自动化指标、专家判断与大模型模拟的重要性,所提流程为中小学语言评估题的开发与验证提供实用且可扩展的方案。

原文摘要 · Abstract (English)

This study explores automatic generation (AIG) using language models to create multiple choice questions (MCQs) for morphological assessment, aiming to reduce the cost and inconsistency of manual test development. The study used a two-fold approach. First, we compared a fine-tuned medium model (Gemma, 2B) with a larger untuned one (GPT-3.5, 175B). Second, we evaluated seven structured prompting strategies, including zero-shot, few-shot, chain-of-thought, role-based, sequential, and combinations. Generated items were assessed using automated metrics and expert scoring across five dimensions. We also used GPT-4.1, trained on expert-rated samples, to simulate human scoring at scale. Results show that structured prompting, especially strategies combining chain-of-thought and sequential design, significantly improved Gemma's outputs. Gemma generally produced more construct-aligned and instructionally appropriate items than GPT-3.5's zero-shot responses, with prompt design playing a key role in mid-size model performance. This study demonstrates that structured prompting and efficient fine-tuning can enhance midsized models for AIG under limited data conditions. We highlight the value of combining automated metrics, expert judgment, and large-model simulation to ensure alignment with assessment goals. The proposed workflow offers a practical and scalable way to develop and validate language assessment items for K-12.

自动出题提示工程小模型教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。