arXiv:2510.19866cs.CLcs.AI2025-10

对比五种AI模型与三种提示框架,找出生成高中物理教案的最佳组合。

An Evaluation of the Pedagogical Soundness and Usability of AI-Generated Lesson Plans Across Different Models and Prompt Frameworks in High-School Physics

  • 用TAG、RACE、COSTAR三种提示框架,测试不同AI模型生成教案效果。
  • RACE框架减少错误率,提升课程标准对齐度,深思模型最易读(FKGL=8.64)。
  • 多数目标停留在记忆理解层级,需加显式清单提升高阶思维设计。

本研究评估了五种主流大语言模型(ChatGPT/GPT-5、Claude Sonnet 4.5、Gemini 2.5 Flash、DeepSeek V3.2、Grok 4)在三种结构化提示框架(TAG、RACE、COSTAR)下生成的高中物理《电磁波谱》教案的教育合理性和可用性。共生成15份教案,通过四项自动化指标分析:(1)可读性与语言复杂度(FKGL),(2)事实准确性与幻觉检测,(3)课程标准与教学大纲对齐度,(4)学习目标的认知要求。结果显示,模型选择显著影响可读性:DeepSeek生成的教案最易读(FKGL=8.64),Claude语言最密集(FKGL=19.89)。提示框架对准确性和教学完整性影响最大,其中RACE框架幻觉指数最低,且与NGSS课程标准的附带对齐度最高。所有教案的学习目标主要集中在布卢姆分类法的“记忆”与“理解”层级,高阶动词使用有限。研究建议:最优配置为选用可读性优的模型,结合RACE框架,并添加物理概念、课程标准和高阶目标的显式检查清单。

原文摘要 · Abstract (English)

This study evaluates the pedagogical soundness and usability of AI-generated lesson plans across five leading large language models: ChatGPT (GPT-5), Claude Sonnet 4.5, Gemini 2.5 Flash, DeepSeek V3.2, and Grok 4. Beyond model choice, three structured prompt frameworks were tested: TAG (Task, Audience, Goal), RACE (Role, Audience, Context, Execution), and COSTAR (Context, Objective, Style, Tone, Audience, Response Format). Fifteen lesson plans were generated for a single high-school physics topic, The Electromagnetic Spectrum. The lesson plans were analyzed through four automated computational metrics: (1) readability and linguistic complexity, (2) factual accuracy and hallucination detection, (3) standards and curriculum alignment, and (4) cognitive demand of learning objectives. Results indicate that model selection exerted the strongest influence on linguistic accessibility, with DeepSeek producing the most readable teaching plan (FKGL = 8.64) and Claude generating the densest language (FKGL = 19.89). The prompt framework structure most strongly affected the factual accuracy and pedagogical completeness, with the RACE framework yielding the lowest hallucination index and the highest incidental alignment with NGSS curriculum standards. Across all models, the learning objectives in the fifteen lesson plans clustered at the Remember and Understand tiers of Bloom's taxonomy. There were limited higher-order verbs in the learning objectives extracted. Overall, the findings suggest that readability is significantly governed by model design, while instructional reliability and curricular alignment depend more on the prompt framework. The most effective configuration for lesson plans identified in the results was to combine a readability-optimized model with the RACE framework and an explicit checklist of physics concepts, curriculum standards, and higher-order objectives.

AI教育提示工程教学设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。