首个面向真实场景的孟加拉语文字识别基准,揭示大模型不必然更优。
BanglaWild: An In-the-Wild Bengali Scene Text Recognition Benchmark for OCR and Vision-Language Models

- 构建2535张真实场景图像与标准转录数据集,含多维诊断属性。
- 发现最强系统仍有60%错误源于视觉误识别,连体字错误不足2%。
- 适合关注低资源语言、视觉语言模型或OCR鲁棒性研究者。
真实场景下的孟加拉语文字识别长期缺乏评估标准:现有资源集中于手写文档或受控标识牌解析,仅报告整体编辑距离,且仅评估传统OCR或视觉语言模型(VLM),从未在相同真实数据上同时评测两者。为填补空白,我们提出BANGLAWILD,一个包含2,535张孟加拉语场景文本图像的数据集,每张图像配有逐字真实转录、两个分类轴、四个诊断属性,以及一个标准化形式(当图像中文字偏离标准拼写时)。我们在三种提示策略下评估15个VLM和3个传统OCR系统,对6个开源模型使用LoRA微调,并引入大模型作为评判者补充编辑距离评估。结果表明,同一模型族中大模型并不优于小模型;15类错误分类显示,最强系统约60%错误源于视觉误识别,而连体字相关错误占比不足2%,挑战了孟加拉语OCR领域长期假设;该视觉主导模式在所有架构中一致,包括唯一能可靠读取孟加拉语的传统基线。提示语言主要影响跨脚本漂移,而LoRA可减少弱模型的灾难性失败,但无法提升已表现良好的模型上限。代码与数据将公开发布。
原文摘要 · Abstract (English)
In-the-wild Bengali scene text recognition is largely unmeasured: existing resources target handwritten documents or constrained sign-board parsing, report only aggregate edit-distance metrics, and evaluate either conventional OCR or VLMs, never both on the same in-the-wild data. To address this gap, we introduce BANGLAWILD, a benchmark of 2,535 Bengali scene text images, each paired with a verbatim gold transcription, two categorical axes, four diagnostic attributes, and an orthographically standard form where the in-image text deviates from canonical spelling. We evaluate fifteen VLMs and three conventional OCR systems under three prompting strategies, fine-tune 6 open-source models with LoRA, and complement edit-distance metrics with an LLM-as-a-Judge evaluation. Our results reveal a persistent gap in which larger models within the same family do not outperform smaller ones. Our fifteen-class error taxonomy shows that visual mis-recognition accounts for ~60% of errors in the strongest systems, while conjunct-related errors contribute under 2%, challenging a long-standing assumption in Bengali OCR research; the same visual dominant profile also holds across architectures, including the one conventional baseline that reads Bengali reliably. Prompt language mainly affects cross-script drift and LoRA reduces catastrophic failures in weak models without lifting the ceiling on already competent ones. Code and data will be publicly released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。