arXiv:2602.16430cs.CVcs.AI2026-02被引 1

为印度多语言文档设计高效高精度的生产级OCR系统

Designing Production-Scale OCR for India: Multilingual and Domain-Specific Systems

  • 用预训练OCR模型微调,比端到端视觉语言模型更优
  • Chitrapathak-2速度提升3-6倍,泰卢固语准确率领先
  • 专用于9类政府文件,结构化字段提取准确率达89.8%

为印度设计生产级OCR系统需兼顾语言多样性、文档异构性与部署限制。本文通过Chitrapathak系列研究两种多语言OCR训练策略:一是将通用视觉编码器与强多语言语言模型端到端联合训练;二是对未针对目标语言训练的现有OCR模型进行微调。在多语言印地语OCR基准和面向部署的指标上评估发现,第二种策略始终具备更优的准确率-延迟权衡。Chitrapathak-2相比前代实现3-6倍加速,在泰卢固语上达SOTA(6.69 char ANLS),其余语言居第二。此外,我们提出Parichay系列独立模型,专为9类印度政府文件设计,可提取结构化关键字段,精确匹配率达89.8%,推理更快。两者共同达成SOTA性能,并为构建印度场景下的生产级OCR流水线提供实用指导。

原文摘要 · Abstract (English)

Designing Optical Character Recognition (OCR) systems for India requires balancing linguistic diversity, document heterogeneity, and deployment constraints. In this paper, we study two training strategies for building multilingual OCR systems with Vision-Language Models through the Chitrapathak series. We first follow a popular multimodal approach, pairing a generic vision encoder with a strong multilingual language model and training the system end-to-end for OCR. Alternatively, we explore fine-tuning an existing OCR model, despite not being trained for the target languages. Through extensive evaluation on multilingual Indic OCR benchmarks and deployment-oriented metrics, we find that the second strategy consistently achieves better accuracy-latency trade-offs. Chitrapathak-2 achieves 3-6x speedup over its predecessor with being state-of-the-art (SOTA) in Telugu (6.69 char ANLS) and second best in the rest. In addition, we present Parichay, an independent OCR model series designed specifically for 9 Indian government documents to extract structured key fields, achieving 89.8% Exact Match score with a faster inference. Together, these systems achieve SOTA performance and provide practical guidance for building production-scale OCR pipelines in the Indian context.

OCR多语言印度生产系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。