构建首个覆盖11种印度语的场景文本数据集,推动多语言文字识别研究。
Bharat Scene Text: A Novel Comprehensive Dataset and Benchmark for Indian Language Scene Text Understanding
- 构建包含10万+词汇的跨语言场景文本数据集,覆盖11种印度语言
- 在6500+图像上标注,支持检测、识别、脚本分类等四项任务
- 开源模型与数据,助力无障碍技术与电商应用研究
场景文本识别在辅助技术、搜索和电商等领域有广泛应用。尽管英语场景文本识别已接近成熟,但印度语言场景文本识别仍面临挑战,主要源于文字多样、字体不规范及缺乏高质量数据集和开源模型。为此,我们提出印度场景文本数据集(Bharat Scene Text Dataset, BSTD),涵盖超过10万词汇,覆盖11种印度语言及英语,来源于6500多张来自印度各地的场景图像。数据集经过精细标注,支持多项任务:(i)场景文本检测,(ii)脚本识别,(iii)裁剪词识别,(iv)端到端场景文本识别。我们对原为英文设计的先进模型进行微调评估,揭示了印度语言识别的难点与机遇。所有模型与数据均开源,有望推动该领域研究进展。
原文摘要 · Abstract (English)
Reading scene text, that is, text appearing in images, has numerous application areas, including assistive technology, search, and e-commerce. Although scene text recognition in English has advanced significantly and is often considered nearly a solved problem, Indian language scene text recognition remains an open challenge. This is due to script diversity, non-standard fonts, and varying writing styles, and, more importantly, the lack of high-quality datasets and open-source models. To address these gaps, we introduce the Bharat Scene Text Dataset (BSTD) - a large-scale and comprehensive benchmark for studying Indian Language Scene Text Recognition. It comprises more than 100K words that span 11 Indian languages and English, sourced from over 6,500 scene images captured across various linguistic regions of India. The dataset is meticulously annotated and supports multiple scene text tasks, including: (i) Scene Text Detection, (ii) Script Identification, (iii) Cropped Word Recognition, and (iv) End-to-End Scene Text Recognition. We evaluated state-of-the-art models originally developed for English by adapting (fine-tuning) them for Indian languages. Our results highlight the challenges and opportunities in Indian language scene text recognition. We believe that this dataset represents a significant step toward advancing research in this domain. All our models and data are open source.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。