用迁移学习提升课件视频中图表等视觉元素的自动识别能力。
Visual Content Detection in Educational Videos with Transfer Learning and Dataset Enrichment
- 采用迁移学习+半监督自动标注,优化YOLO模型用于课件视频检测。
- 在多个基准数据集上训练后,检测准确率显著优于通用模型。
- 开源了标注数据集与代码,助力教育视频智能分析研究。
视频正重塑教育,线上课程和录播讲座逐步替代或补充课堂教学。现有研究聚焦于提升视频讲座的信息检索能力,包括高级导航、搜索、摘要生成及问答聊天机器人。视觉元素如表格、图表和插图对理解、记忆和数据呈现至关重要,但其在提升视频内容可访问性方面的潜力尚未被充分挖掘。主要原因在于:第一,多数视觉元素(如图表、图形、表格、插图)为人工创建,缺乏标准结构;第二,连贯的视觉对象可能边界模糊,由文本与图像组件共同构成。尽管深度学习目标检测技术进步显著,但现有模型在课件视频中的表现仍不理想,受限于内容独特性及标注数据稀缺。本文提出一种基于迁移学习的视觉元素检测方法,评估了多种先进目标检测模型在课件视频数据集上的表现,发现YOLO最具潜力。随后通过多数据集训练并引入半监督自动标注策略优化该模型。实验验证了该方法的有效性,为课件视频中的目标检测提供通用解决方案。论文贡献包括一个公开发布的标注课件视频帧数据集及源代码,以促进后续研究。
原文摘要 · Abstract (English)
Video is transforming education with online courses and recorded lectures supplementing and replacing classroom teaching. Recent research has focused on enhancing information retrieval for video lectures with advanced navigation, searchability, summarization, as well as question answering chatbots. Visual elements like tables, charts, and illustrations are central to comprehension, retention, and data presentation in lecture videos, yet their full potential for improving access to video content remains underutilized. A major factor is that accurate automatic detection of visual elements in a lecture video is challenging; reasons include i) most visual elements, such as charts, graphs, tables, and illustrations, are artificially created and lack any standard structure, and ii) coherent visual objects may lack clear boundaries and may be composed of connected text and visual components. Despite advancements in deep learning based object detection, current models do not yield satisfactory performance due to the unique nature of visual content in lectures and scarcity of annotated datasets. This paper reports on a transfer learning approach for detecting visual elements in lecture video frames. A suite of state of the art object detection models were evaluated for their performance on lecture video datasets. YOLO emerged as the most promising model for this task. Subsequently YOLO was optimized for lecture video object detection with training on multiple benchmark datasets and deploying a semi-supervised auto labeling strategy. Results evaluate the success of this approach, also in developing a general solution to the problem of object detection in lecture videos. Paper contributions include a publicly released benchmark of annotated lecture video frames, along with the source code to facilitate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。