用对话式AI生成更深入的教学反思题,提升学生思考质量
Reflecting in the Reflection: Integrating a Socratic Questioning Framework into Automated AI-Based Question Generation
- 双角色AI协作:学生教师提问题,教育教师逐轮优化
- 动态对话终止机制使问题相关性和深度显著优于固定流程
- 适合教育科技研究者和一线教师开发智能教学工具
设计高质量的反思性问题在教学中至关重要,但耗时且教师支持不均。本文提出一种‘反思中的反思’框架,利用大语言模型(LLM)自动生成反思问题。该框架由两名角色专用代理——学生教师与教育教师——协同工作,通过苏格拉底式多轮对话,基于教师指定的主题、核心概念、学生水平及可选教学材料,迭代优化单一问题。学生教师提出候选问题并附简要理由,教育教师则从清晰度、深度、相关性、参与度和概念关联性五方面评估,并仅以针对性引导问题或终止信号回应。实验在真实初中信息技术课堂中进行,以GPT-4o-mini为骨干模型,使用更强的GPT-4级模型作为外部评价者,在成对比较中评估清晰度、相关性、深度与整体质量。结果表明:动态停止策略结合上下文信息(如学生水平、材料)显著优于固定5步或10步迭代;多轮对话下问题整体质量更高,相关性与深度显著提升,优于单次生成基线。
原文摘要 · Abstract (English)
Designing good reflection questions is pedagogically important but time-consuming and unevenly supported across teachers. This paper introduces a reflection-in-reflection framework for automated generation of reflection questions with large language models (LLMs). Our approach coordinates two role-specialized agents, a Student-Teacher and a Teacher-Educator, that engage in a Socratic multi-turn dialogue to iteratively refine a single question given a teacher-specified topic, key concepts, student level, and optional instructional materials. The Student-Teacher proposes candidate questions with brief rationales, while the Teacher-Educator evaluates them along clarity, depth, relevance, engagement, and conceptual interconnections, responding only with targeted coaching questions or a fixed signal to stop the dialogue. We evaluate the framework in an authentic lower-secondary ICT setting on the topic, using GPT-4o-mini as the backbone model and a stronger GPT- 4-class LLM as an external evaluator in pairwise comparisons of clarity, relevance, depth, and overall quality. First, we study how interaction design and context (dynamic vs.fixed iteration counts; presence or absence of student level and materials) affect question quality. Dynamic stopping combined with contextual information consistently outperforms fixed 5- or 10-step refinement, with very long dialogues prone to drift or over-complication. Second, we show that our two-agent protocol produces questions that are judged substantially more relevant and deeper, and better overall, than a one-shot baseline using the same backbone model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。