arXiv:2509.17961cs.CL2025-09EMNLP被引 2

为虚拟助教问答能力设计基于教育科学的评估框架

Bringing Pedagogy into Focus: Evaluating Virtual Teaching Assistants' Question-Answering in Asynchronous Learning Environments

  • 基于学习科学构建针对异步论坛的评估框架
  • 通过专家标注训练分类器,准确识别问答质量
  • 助力教育AI实现更符合教学规律的智能支持

异步学习环境(ALE)广泛应用于正式与非正式学习,但及时且个性化的支持常显不足。在此背景下,虚拟教学助理(VTA)可减轻教师负担,但其评估需兼具严谨性与教育理论基础。现有评估多依赖表面指标,缺乏教育学支撑,难以有效比较不同VTA系统的教学效果。为此,本文提出一个根植于学习科学、专用于异步论坛讨论的评估框架。我们利用专家对VTA回复的标注数据,构建分类模型,并评估其有效性,识别出提升准确率的方法及阻碍泛化的问题。本研究为VTA系统提供了理论驱动的评估基础,推动教育AI向更具教学实效的方向发展。

原文摘要 · Abstract (English)

Asynchronous learning environments (ALEs) are widely adopted for formal and informal learning, but timely and personalized support is often limited. In this context, Virtual Teaching Assistants (VTAs) can potentially reduce the workload of instructors, but rigorous and pedagogically sound evaluation is essential. Existing assessments often rely on surface-level metrics and lack sufficient grounding in educational theories, making it difficult to meaningfully compare the pedagogical effectiveness of different VTA systems. To bridge this gap, we propose an evaluation framework rooted in learning sciences and tailored to asynchronous forum discussions, a common VTA deployment context in ALE. We construct classifiers using expert annotations of VTA responses on a diverse set of forum posts. We evaluate the effectiveness of our classifiers, identifying approaches that improve accuracy as well as challenges that hinder generalization. Our work establishes a foundation for theory-driven evaluation of VTA systems, paving the way for more pedagogically effective AI in education.

教育AI虚拟助教学习科学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。