开源模型Typhoon OCR精准提取泰语文档,性能媲美大厂闭源系统。
Typhoon OCR: Open Vision-Language Model For Thai Document Extraction
- 基于多阶段数据构建,微调中文本语言模型支持泰语
- 在金融报告等5类文档上表现超主流闭源模型,计算成本低
- 轻量高效,无需依赖元数据,适合实际部署
文档提取是数字工作流的核心,但现有视觉语言模型(VLM)主要面向高资源语言。泰语因非拉丁字母、无明确词边界及真实场景文档高度无序,给现有开源模型带来挑战。本文提出针对泰语和英语的开源视觉语言模型Typhoon OCR。该模型基于视觉语言主干网络,使用泰语专用训练数据进行微调。数据集通过多阶段构建流程生成:结合传统OCR、VLM重构与人工校准的合成数据。Typhoon OCR实现统一框架下的文本转录、版面重建与文档结构一致性保持。最新版本V1.5为紧凑型轻量模型,降低对元数据依赖,提升部署便利性。在财务报告、政府表格、书籍、信息图及手写文档等五类泰语文档上的综合评估显示,其性能可比肩甚至超越更大规模的前沿闭源模型,同时计算开销显著更低。结果表明,开源视觉语言文档提取模型可在泰语上实现高精度文本与版面重建,达到闭源系统水平,且更轻量易部署。
原文摘要 · Abstract (English)
Document extraction is a core component of digital workflows, yet existing vision-language models (VLMs) predominantly favor high-resource languages. Thai presents additional challenges due to script complexity from non-latin letters, the absence of explicit word boundaries, and the prevalence of highly unstructured real-world documents, limiting the effectiveness of current open-source models. This paper presents Typhoon OCR, an open VLM for document extraction tailored for Thai and English. The model is fine-tuned from vision-language backbones using a Thai-focused training dataset. The dataset is developed using a multi-stage data construction pipeline that combines traditional OCR, VLM-based restructuring, and curated synthetic data. Typhoon OCR is a unified framework capable of text transcription, layout reconstruction, and document-level structural consistency. The latest iteration of our model, Typhoon OCR V1.5, is a compact and inference-efficient model designed to reduce reliance on metadata and simplify deployment. Comprehensive evaluations across diverse Thai document categories, including financial reports, government forms, books, infographics, and handwritten documents, show that Typhoon OCR achieves performance comparable to or exceeding larger frontier proprietary models, despite substantially lower computational cost. The results demonstrate that open vision-language OCR models can achieve accurate text extraction and layout reconstruction for Thai documents, reaching performance comparable to proprietary systems while remaining lightweight and deployable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。