构建1.2万张代码录屏图像数据集,助力编程教学视频分析
CodeSCAN: ScreenCast ANalysis for Video Programming Tutorials
- 从VS Code环境采集1.2万张截图,覆盖24种语言与90+主题
- 包含真实用户操作与多变布局,支持IDE元素检测与OCR评测
- 开源数据集和评估工具,适合做视频编程教育研究者使用
以编码录屏形式的编程教程在编程教育中至关重要,服务新手与资深开发者。然而视频格式带来搜索困难。针对现有录屏分析缺乏大规模多样数据集的问题,我们提出CodeSCAN数据集,包含12,000张在Visual Studio Code开发环境中截取的屏幕图像,涵盖24种编程语言、25种字体及超过90种不同主题,同时包含多样的布局变化和真实的用户交互行为。此外,我们对IDE元素检测、颜色转灰度、光学字符识别(OCR)进行了详尽的定量与定性评估。我们希望本工作能推动编码录屏分析领域的研究,并将数据集生成代码和基准测试工具公开于官网。
原文摘要 · Abstract (English)
Programming tutorials in the form of coding screencasts play a crucial role in programming education, serving both novices and experienced developers. However, the video format of these tutorials presents a challenge due to the difficulty of searching for and within videos. Addressing the absence of large-scale and diverse datasets for screencast analysis, we introduce the CodeSCAN dataset. It comprises 12,000 screenshots captured from the Visual Studio Code environment during development, featuring 24 programming languages, 25 fonts, and over 90 distinct themes, in addition to diverse layout changes and realistic user interactions. Moreover, we conduct detailed quantitative and qualitative evaluations to benchmark the performance of Integrated Development Environment (IDE) element detection, color-to-black-and-white conversion, and Optical Character Recognition (OCR). We hope that our contributions facilitate more research in coding screencast analysis, and we make the source code for creating the dataset and the benchmark publicly available on this website.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。