用比赛式评估法找出最有效的教育类提示,提升个性化学习效果。
LLM Prompt Evaluation for Educational Applications
- 设计6种提示模板,融合不同教学策略进行对比测试。
- 一种聚焦策略性阅读的提示胜率高达81%~100%。
- 适合教育科技研究者系统优化AI教学提示,推动实证驱动设计。
随着大语言模型在教育应用中日益普及,亟需基于证据的方法来设计和评估生成个性化、符合教学目标输出的提示。本研究提出一种可推广的系统化评估方法,通过分析结构化对话活动中模型生成的追问问题展开。设计了六种提示模板,融合已知提示工程模式,每种强调不同教学策略。采用锦标赛式评估框架,使用Glicko2评分系统,由八位评审员从格式、对话支持度、学习者适切性三个维度评估问题对。数据来自三个不同教育场景下的120次真实用户交互。结果显示,一种与策略性阅读相关的提示在两两比较中胜率达81%至100%,该提示结合角色设定与上下文管理模式,旨在支持元认知学习策略如自主学习。该方法为教育技术研究者提供了系统评估与改进提示设计的路径,推动从随意提示工程迈向基于证据的提示开发。
原文摘要 · Abstract (English)
As large language models (LLMs) become increasingly common in educational applications, there is a growing need for evidence-based methods to design and evaluate LLM prompts that produce personalized and pedagogically aligned out-puts. This study presents a generalizable, systematic approach for evaluating prompts, demonstrated through an analysis of LLM-generated follow-up questions in a structured dialogue activity. Six prompt templates were designed and tested. The templates incorporated established prompt engineering patterns, with each prompt emphasizing distinct pedagogical strategies. The prompt templates were compared through a tournament-style evaluation framework that can be adapted for other educational applications. The tournament employed the Glicko2 rating system with eight judges evaluating question pairs across three dimensions: format, dialogue support, and appropriateness for learners. Data was sourced from 120 authentic user interactions across three distinct educational deployments. Results showed that a single prompt related to strategic reading out-performed other templates with win probabilities ranging from 81% to 100% in pairwise comparisons. This prompt combined persona and context manager pat-terns and was designed to support metacognitive learning strategies such as self-directed learning. The methodology showcases how educational technology re- searchers can systematically evaluate and improve prompt designs, moving beyond ad-hoc prompt engineering toward evidence-based prompt development for educational applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。