arXiv:2505.01790cs.CVcs.CL2025-05被引 4

用视觉语言模型自动生成教育视频问题,提升学习参与度。

Enhancing the Learning Experience: Using Vision-Language Models to Generate Questions for Educational Videos

  • 利用视觉语言模型分析视频内容生成学习问题
  • 微调后问题相关性与可答性显著提升
  • 适合教育科技研究者与在线课程设计者参考

基于网络的教育视频提供了灵活的学习机会,日益受到欢迎。然而,如何提高用户参与度和知识留存仍是挑战。自动生的问题能激发学习者并支持知识获取,还能帮助教师和学习者评估理解程度。尽管大语言模型和视觉语言模型已在多种任务中应用,但其在教育视频问答生成中的应用仍不充分。本文研究当前视觉语言模型生成教育视频学习问题的能力,评估(1)零样本模型表现;(2)微调对内容特定问题生成的影响;(3)不同视频模态对问题质量的影响;(4)通过定性研究分析生成问题的相关性、可答性和难度水平。研究揭示了现有模型的潜力,强调了微调的必要性,并指出问题多样性与相关性的挑战。文中提出了未来多模态数据集的需求及有前景的研究方向。

原文摘要 · Abstract (English)

Web-based educational videos offer flexible learning opportunities and are becoming increasingly popular. However, improving user engagement and knowledge retention remains a challenge. Automatically generated questions can activate learners and support their knowledge acquisition. Further, they can help teachers and learners assess their understanding. While large language and vision-language models have been employed in various tasks, their application to question generation for educational videos remains underexplored. In this paper, we investigate the capabilities of current vision-language models for generating learning-oriented questions for educational video content. We assess (1) out-of-the-box models' performance; (2) fine-tuning effects on content-specific question generation; (3) the impact of different video modalities on question quality; and (4) in a qualitative study, question relevance, answerability, and difficulty levels of generated questions. Our findings delineate the capabilities of current vision-language models, highlighting the need for fine-tuning and addressing challenges in question diversity and relevance. We identify requirements for future multimodal datasets and outline promising research directions.

教育AI视觉语言自动问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。