arXiv:2411.05043cs.CVcs.AI2024-11

构建多语言视频字幕数据集,支持跨语言文本识别研究

Multi-language Video Subtitle Dataset for Image-based Text Recognition

  • 从24段视频提取4224张字幕图像,涵盖157种字符
  • 包含泰语、罗马字母、阿拉伯数字等复杂字符与布局
  • 适合研究视频中复杂背景下的文本识别模型

多语言视频字幕数据集是一个为多语言文本识别研究设计的综合性资源,包含从24段网络视频中提取的4224张字幕图像。数据集涵盖泰语辅音、元音、声调标记、标点符号、数字、罗马字母及阿拉伯数字等多种字符,共157个唯一字符。图像中文字长度、字体和位置差异显著,模拟真实复杂背景。该数据集响应了视频嵌入字幕日益普及带来的高质量多语言文本识别数据需求,尤其在YouTube、Facebook等平台背景下具有重要价值。它为训练与评估深度学习模型提供可靠基准,助力提升文本识别系统的准确性与计算效率,推动人工智能、深度学习、计算机视觉与模式识别等领域的发展。

原文摘要 · Abstract (English)

The Multi-language Video Subtitle Dataset is a comprehensive collection designed to support research in text recognition across multiple languages. This dataset includes 4,224 subtitle images extracted from 24 videos sourced from online platforms. It features a wide variety of characters, including Thai consonants, vowels, tone marks, punctuation marks, numerals, Roman characters, and Arabic numerals. With 157 unique characters, the dataset provides a resource for addressing challenges in text recognition within complex backgrounds. It addresses the growing need for high-quality, multilingual text recognition data, particularly as videos with embedded subtitles become increasingly dominant on platforms like YouTube and Facebook. The variability in text length, font, and placement within these images adds complexity, offering a valuable resource for developing and evaluating deep learning models. The dataset facilitates accurate text transcription from video content while providing a foundation for improving computational efficiency in text recognition systems. As a result, it holds significant potential to drive advancements in research and innovation across various computer science disciplines, including artificial intelligence, deep learning, computer vision, and pattern recognition.

视频字幕多语言识别文本识别数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。