系统梳理大模型提示工程在中小学科学教育中的应用与效果
A Systematic Review on Prompt Engineering in Large Language Models for K-12 STEM Education
- 按PRISMA标准筛选30篇2021–2024年研究,分析提示策略与模型使用
- 少样本和思维链提示在教学任务中表现更优,小模型配好提示胜过大模型
- 揭示评估方法不统一、实证验证不足等局限,适合教育AI研究者参考
大型语言模型(LLMs)有潜力通过提升教学与学习过程来改善中小学科学教育。尽管已有研究显示积极成果,但对如何通过提示工程有效应用LLMs仍缺乏全面理解。为此,本研究基于PRISMA协议,筛选了2,654篇文献,最终纳入30项实证研究进行分析。结果表明,虽然简单提示和零样本提示仍占主导,但少样本和思维链提示在各类教育任务中表现出更优效果。主流使用GPT系列模型,但在特定场景下,小型或微调模型(如Blender 7B)配合有效提示优于大模型(如GPT-3)。评估方法差异显著,且多数研究缺乏真实环境下的实证验证。
原文摘要 · Abstract (English)
Large language models (LLMs) have the potential to enhance K-12 STEM education by improving both teaching and learning processes. While previous studies have shown promising results, there is still a lack of comprehensive understanding regarding how LLMs are effectively applied, specifically through prompt engineering-the process of designing prompts to generate desired outputs. To address this gap, our study investigates empirical research published between 2021 and 2024 that explores the use of LLMs combined with prompt engineering in K-12 STEM education. Following the PRISMA protocol, we screened 2,654 papers and selected 30 studies for analysis. Our review identifies the prompting strategies employed, the types of LLMs used, methods of evaluating effectiveness, and limitations in prior work. Results indicate that while simple and zero-shot prompting are commonly used, more advanced techniques like few-shot and chain-of-thought prompting have demonstrated positive outcomes for various educational tasks. GPT-series models are predominantly used, but smaller and fine-tuned models (e.g., Blender 7B) paired with effective prompt engineering outperform prompting larger models (e.g., GPT-3) in specific contexts. Evaluation methods vary significantly, with limited empirical validation in real-world settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。