arXiv:2505.24477cs.CYcs.AI2025-05被引 19

用真人教师盲测对比,发现Gemini 2.5 Pro最懂教学。

Evaluating Gemini in an arena for learning

  • 让189位教师模拟真实教学场景,对比AI模型表现
  • 专家评判中73.2%更偏好Gemini 2.5 Pro,排名第一
  • 在教学原理契合度上显著优于其他模型,适合教育应用

人工智能有望重塑教育,但研究界缺乏可靠的通用评估基准。为评测顶尖AI在教育场景中的表现,我们开展了一场“学习竞技场”,由教育工作者与教学法专家对主流AI模型进行盲评、多轮、一对一比较。共189位教育者基于经验扮演真实学习场景,先后与两个模型交互,随后206位专家判断哪个模型更有效支持学习目标。评估涵盖Gemini 2.5 Pro、Claude 3.7 Sonnet、GPT-4o和OpenAI o3等前沿模型。排除平局后,专家在73.2%的对比中更青睐Gemini 2.5 Pro,综合排名首位。该模型在良好教学原则的体现上也显著更优。结果表明,Gemini 2.5 Pro是当前最适配学习场景的AI模型。

原文摘要 · Abstract (English)

Artificial intelligence (AI) is poised to transform education, but the research community lacks a robust, general benchmark to evaluate AI models for learning. To assess state-of-the-art support for educational use cases, we ran an "arena for learning" where educators and pedagogy experts conduct blind, head-to-head, multi-turn comparisons of leading AI models. In particular, $N = 189$ educators drew from their experience to role-play realistic learning use cases, interacting with two models sequentially, after which $N = 206$ experts judged which model better supported the user's learning goals. The arena evaluated a slate of state-of-the-art models: Gemini 2.5 Pro, Claude 3.7 Sonnet, GPT-4o, and OpenAI o3. Excluding ties, experts preferred Gemini 2.5 Pro in 73.2% of these match-ups -- ranking it first overall in the arena. Gemini 2.5 Pro also demonstrated markedly higher performance across key principles of good pedagogy. Altogether, these results position Gemini 2.5 Pro as a leading model for learning.

教育AI模型评测Gemini

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。