arXiv:2603.02775cs.CLcs.LG2026-03AAAI被引 2

评测大模型数学教学能力,发现其解题强但教学弱。

From Solver to Tutor: Evaluating the Pedagogical Intelligence of LLMs with KMP-Bench

  • 构建双模块评测基准,覆盖对话与技能两方面教学能力。
  • 顶尖模型解题准确率高,但教学原则应用能力差。
  • 提供15万条教学对话数据集,提升模型教学效果。

大型语言模型在人工智能数学辅导中展现出巨大潜力,但现有评估多依赖简单指标或狭窄教学场景,难以全面衡量多轮教学的有效性。本文提出KMP-Bench,一个面向K-8年级数学教学的综合性评测基准,包含两个互补模块:KMP-Dialogue通过整合多样教学要素构建的多轮对话数据集,评估模型在挑战、解释、反馈等六大核心教学原则上的综合能力;KMP-Skills则对多轮解题、错误识别与修正、题目生成等基础教学能力进行细粒度评估。在该基准上的评估显示,尽管领先模型在可验证解法任务上表现优异,但在教学原则的细致应用上仍显不足。此外,我们还发布了包含15万条对话的KMP-Pile大规模数据集。在该数据集上微调的模型在KMP-Bench上表现出显著提升,证明了富含教学语境的数据对打造更高效AI数学导师的重要性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) show significant potential in AI mathematical tutoring, yet current evaluations often rely on simplistic metrics or narrow pedagogical scenarios, failing to assess comprehensive, multi-turn teaching effectiveness. In this paper, we introduce KMP-Bench, a comprehensive K-8 Mathematical Pedagogical Benchmark designed to assess LLMs from two complementary perspectives. The first module, KMP-Dialogue, evaluates holistic pedagogical capabilities against six core principles (e.g., Challenge, Explanation, Feedback), leveraging a novel multi-turn dialogue dataset constructed by weaving together diverse pedagogical components. The second module, KMP-Skills, provides a granular assessment of foundational tutoring abilities, including multi-turn problem-solving, error detection and correction, and problem generation. Our evaluations on KMP-Bench reveal a key disparity: while leading LLMs excel at tasks with verifiable solutions, they struggle with the nuanced application of pedagogical principles. Additionally, we present KMP-Pile, a large-scale (150K) dialogue dataset. Models fine-tuned on KMP-Pile show substantial improvement on KMP-Bench, underscoring the value of pedagogically-rich training data for developing more effective AI math tutors.

数学教育大模型评测教学智能对话数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。