首个高棉语场景文本检测识别数据集,助力低资源语言研究
KhmerST: A Low-Resource Khmer Scene Text Detection and Recognition Benchmark
- 构建首个高棉语场景文本数据集,含1544张专家标注图像
- 覆盖室内外、模糊、遮挡等多种复杂场景,支持文本检测与识别
- 提供基准模型,适合低资源语言视觉任务研究者使用
构建高效场景文本检测与识别模型依赖大量训练数据,而获取这些数据对低资源语言而言既耗时又昂贵。传统针对拉丁字符的方法在非拉丁文字中表现不佳,因存在字符堆叠、变音符号及无明确词边界的可变字符宽度等问题。本文首次提出高棉语场景文本数据集(KhmerST),包含1544张经专家标注的图像,涵盖997张室内与547张室外场景。数据集涵盖平面文本、凸起文本、光照不足、远距离及部分遮挡文本。每幅图像均提供行级文本内容及多边形边界框坐标。同时提供场景文本检测与识别任务的基准模型,为后续研究提供坚实起点。该数据集已公开于 https://gitlab.com/vannkinhnom123/khmerst。
原文摘要 · Abstract (English)
Developing effective scene text detection and recognition models hinges on extensive training data, which can be both laborious and costly to obtain, especially for low-resourced languages. Conventional methods tailored for Latin characters often falter with non-Latin scripts due to challenges like character stacking, diacritics, and variable character widths without clear word boundaries. In this paper, we introduce the first Khmer scene-text dataset, featuring 1,544 expert-annotated images, including 997 indoor and 547 outdoor scenes. This diverse dataset includes flat text, raised text, poorly illuminated text, distant and partially obscured text. Annotations provide line-level text and polygonal bounding box coordinates for each scene. The benchmark includes baseline models for scene-text detection and recognition tasks, providing a robust starting point for future research endeavors. The KhmerST dataset is publicly accessible at https://gitlab.com/vannkinhnom123/khmerst.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。