arXiv:2509.24120cs.CL2025-09EMNLP被引 3

用大模型自动回答视频课学生问题,提升在线学习互动性

EduVidQA: Generating and Evaluating Long-form Answers to Student Questions based on Lecture Videos

  • 基于视频内容生成长文本答案,融合视觉与语言信息
  • 构建5252对问答数据集,覆盖296个计算机课程视频
  • 实证研究学生偏好,为教育AI提供真实反馈依据

随着数字平台重塑教育模式,保持互动性对有效学习至关重要。本文探索利用多模态大模型(MLLMs)自动回应在线讲座中的学生提问——这一具有现实意义的新问答任务。我们提出了EduVidQA数据集,包含5252对问答对(合成与真实数据),来自296个涵盖多种主题和难度级别的计算机科学视频。为理解数据集需求与评估标准,我们通过实证研究学生定性偏好,为该领域研究提供重要贡献。基准实验包含6个最先进的MLLMs,评估了合成数据微调的有效性,并揭示了该任务的挑战性。我们采用文本与定性双重指标评估模型表现,展现性能的多维视角,对未来研究至关重要。本工作不仅为该关键问题设立了基准,还为自然语言处理在教育中的应用开辟了新方向。

原文摘要 · Abstract (English)

As digital platforms redefine educational paradigms, ensuring interactivity remains vital for effective learning. This paper explores using Multimodal Large Language Models (MLLMs) to automatically respond to student questions from online lectures - a novel question answering task of real world significance. We introduce the EduVidQA Dataset with 5252 question-answer pairs (both synthetic and real-world) from 296 computer science videos covering diverse topics and difficulty levels. To understand the needs of the dataset and task evaluation, we empirically study the qualitative preferences of students, which we provide as an important contribution to this line of work. Our benchmarking experiments consist of 6 state-of-the-art MLLMs, through which we study the effectiveness of our synthetic data for finetuning, as well as showing the challenging nature of the task. We evaluate the models using both text-based and qualitative metrics, thus showing a nuanced perspective of the models' performance, which is paramount to future work. This work not only sets a benchmark for this important problem, but also opens exciting avenues for future research in the field of Natural Language Processing for Education.

教育AI多模态问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。