首个面向泰语的多模态理解基准,助力AI读懂泰文文档
ThaiOCRBench: A Task-Diverse Benchmark for Vision-Language Understanding in Thai
- 构建13类任务的泰语图文数据集,含2808个标注样本
- 开源模型在手写文本识别上表现远逊于商用模型
- 揭示语言偏见与结构错配等核心挑战,适合低资源语言研究者
我们提出ThaiOCRBench,首个针对泰语文本密集视觉理解任务的综合性评估基准。尽管多模态模型进展迅速,现有基准仍主要集中于高资源语言,泰语在需理解文档结构的任务中长期被忽视。ThaiOCRBench通过包含2,808个样本的多样化、人工标注数据集填补这一空白,涵盖13个任务类别。我们在零样本设置下评估了多种前沿视觉语言模型(VLMs),涵盖专有与开源系统。结果表明存在显著性能差距,专有模型(如Gemini 2.5 Pro)优于开源模型。尤其在细粒度文本识别和手写内容提取任务中,开源模型表现下降最明显。通过详细错误分析,我们识别出语言偏见、结构不匹配及幻觉内容等关键挑战。ThaiOCRBench为低资源、字形复杂场景下的VLM评估提供标准化框架,并为提升泰语文档理解能力提供可操作洞见。
原文摘要 · Abstract (English)
We present ThaiOCRBench, the first comprehensive benchmark for evaluating vision-language models (VLMs) on Thai text-rich visual understanding tasks. Despite recent progress in multimodal modeling, existing benchmarks predominantly focus on high-resource languages, leaving Thai underrepresented, especially in tasks requiring document structure understanding. ThaiOCRBench addresses this gap by offering a diverse, human-annotated dataset comprising 2,808 samples across 13 task categories. We evaluate a wide range of state-of-the-art VLMs in a zero-shot setting, spanning both proprietary and open-source systems. Results show a significant performance gap, with proprietary models (e.g., Gemini 2.5 Pro) outperforming open-source counterparts. Notably, fine-grained text recognition and handwritten content extraction exhibit the steepest performance drops among open-source models. Through detailed error analysis, we identify key challenges such as language bias, structural mismatch, and hallucinated content. ThaiOCRBench provides a standardized framework for assessing VLMs in low-resource, script-complex settings, and provides actionable insights for improving Thai-language document understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。