arXiv:2509.05346cs.AI2025-09被引 1

用真实学习场景评估大模型教学行为,看清谁更懂学生。

Beyond Benchmarking: Scenario-Based Evaluation of Large Language Models for Personalized Learning

  • 设计真实课后辅导场景,让模型诊断知识盲点并给个性化建议。
  • 多模型对比显示差异显著,部分模型误判率超40%且建议模糊。
  • 适合教育AI研发者、教学工具设计师参考,提升模型教学实用性。

尽管大型语言模型(LLMs)被越来越多地用于支持个性化学习,但对其在真实学习情境中教学行为的理解仍很有限。现有评估多依赖基准分数和整体排名,难以揭示模型如何诊断学生理解程度并生成个性化指导。本研究提出一种基于场景的评估框架,以课后辅导为示例,提供学生对数据结构题目的作答数据,要求多个LLM识别隐含知识点、推断掌握情况并生成改进建议。采用Gemini作为外部评估者,在诊断准确性、教学清晰度、可操作性、误解识别及适切性等维度进行一致性评估。通过布拉德利-特雷西模型拟合成对偏好,得到相对强度估计;结合定性分析与语义可视化,进一步考察反馈结构、诊断深度与建议具体性差异。结果表明,不同LLM在同一学习场景中表现出明显可区分的教学行为,证明场景化评估能提供教育意义明确的模型行为洞察。

原文摘要 · Abstract (English)

While large language models (LLMs) are increasingly being adopted to support personalized learning, there remains limited understanding of how their pedagogical behaviors differ in authentic learning scenarios. Existing evaluation practices often emphasize benchmark scores and overall model rankings, but such approaches usually provide limited insight into how LLMs diagnose student understanding and generate personalized guidance. This study proposes a scenario-based evaluation framework for closely examining LLM behavior in personalized learning support. Using a post-class tutoring setting as an illustrative example, a dataset comprising a student's responses to a set of data structures questions is provided to multiple LLMs. Each model is required to identify the underlying knowledge concepts, infer the student's mastery profile, and generate personalized guidance for improvement. To support consistent, reproducible and scalable comparison, Gemini is employed as an external evaluator across multiple pedagogically relevant dimensions, including diagnostic accuracy, instructional clarity, actionability, misconception identification, and appropriateness to the student's level. The resulting pairwise preferences are then fitted using the Bradley-Terry model to derive comparative strength estimates, while qualitative analysis and semantic visualization are used to further examine differences in feedback structure, diagnostic depth, and recommendation specificity. The key findings show that different LLMs exhibit distinguishable pedagogical behaviors within the same learning scenario and demonstrates how scenario-based evaluation can provide educationally meaningful signals for understanding model behavior in AI-enhanced education.

个性化学习大模型评估教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。