让大模型像老师一样教人,更懂教学逻辑。
Pedagogy-R1: Pedagogically-Aligned Reasoning Model with Balanced Educational Benchmark
- 用教学逻辑优化模型输出,提升可教性
- 构建多维度教育评估基准,覆盖知识与教学行为
- 通过教师式推理提示,生成自然教学流程
近期大型推理模型在数学、编程等结构化领域表现优异,但缺乏教学连贯性和真实教学行为。为此,我们提出Pedagogy-R1框架,包含三项创新:(1) 基于知识蒸馏的流水线,过滤并优化模型输出用于指令微调;(2) 全面教育评估基准(WBEB),从学科知识、教学知识、解题追踪、作文评分到教师决策等多维度评估;(3) 教学链(CoP)提示策略,生成和激发教师风格的推理过程。混合方法评估结合定量指标与定性分析,首次系统评估了大型推理模型在教学能力上的优劣。
原文摘要 · Abstract (English)
Recent advances in large reasoning models (LRMs) show strong performance in structured domains such as mathematics and programming; however, they often lack pedagogical coherence and realistic teaching behaviors. To bridge this gap, we introduce Pedagogy-R1, a framework that adapts LRMs for classroom use through three innovations: (1) a distillation-based pipeline that filters and refines model outputs for instruction-tuning, (2) the Well-balanced Educational Benchmark (WBEB), which evaluates performance across subject knowledge, pedagogical knowledge, tracing, essay scoring, and teacher decision-making, and (3) a Chain-of-Pedagogy (CoP) prompting strategy for generating and eliciting teacher-style reasoning. Our mixed-method evaluation combines quantitative metrics with qualitative analysis, providing the first systematic assessment of LRMs' pedagogical strengths and limitations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。