构建教育视频视觉物体检测新数据集,支持教学内容自动识别
Lecture Video Visual Objects (LVVO) Dataset: A Benchmark for Visual Object Detection in Educational Videos
- 从245段课程视频中提取4000帧,标注四类教学视觉元素
- 1000帧经双人标注+专家仲裁,标注一致率达83.41%
- 采用半监督方法扩展至3000帧,支持多类学习范式研究
我们提出了讲座视频视觉对象(LVVO)数据集,为教育视频中的视觉物体检测提供新基准。该数据集包含从245段涵盖生物学、计算机科学和地球科学的讲座视频中提取的4000帧图像。其中1000帧(称为LVVO_1k)被人工标注了四种视觉类别:表格、图表、照片图像和视觉插图,每帧由两名标注员独立标注,交叠标注的F1分数达83.41%,表明标注一致性较高。为确保高质量共识,所有分歧由第三位专家通过冲突解决流程裁定。其余3000帧则通过半监督方法自动标注,形成LVVO_3k。完整数据集为开发与评估教育视频中视觉内容检测的监督与半监督方法提供了宝贵资源。该数据集已公开,以支持本领域的后续研究。
原文摘要 · Abstract (English)
We introduce the Lecture Video Visual Objects (LVVO) dataset, a new benchmark for visual object detection in educational video content. The dataset consists of 4,000 frames extracted from 245 lecture videos spanning biology, computer science, and geosciences. A subset of 1,000 frames, referred to as LVVO_1k, has been manually annotated with bounding boxes for four visual categories: Table, Chart-Graph, Photographic-image, and Visual-illustration. Each frame was labeled independently by two annotators, resulting in an inter-annotator F1 score of 83.41%, indicating strong agreement. To ensure high-quality consensus annotations, a third expert reviewed and resolved all cases of disagreement through a conflict resolution process. To expand the dataset, a semi-supervised approach was employed to automatically annotate the remaining 3,000 frames, forming LVVO_3k. The complete dataset offers a valuable resource for developing and evaluating both supervised and semi-supervised methods for visual content detection in educational videos. The LVVO dataset is publicly available to support further research in this domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。