arXiv:2509.03615cs.CLcs.AI2025-09综述被引 2

对比5种大模型与2种传统OCR,发现传统系统更适合边缘部署。

E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition

  • 构建专为边缘设备优化的Sprinklr-Edge-OCR系统。
  • 传统OCR处理速度比大模型快35倍,成本不足其1%。
  • 在资源受限场景下,传统方案综合性能最优。

多语言、噪声复杂的真实图像中的光学字符识别(OCR)仍是重大挑战。随着大型视觉语言模型(LVLMs)兴起,其泛化与推理能力备受关注。本文提出专为边缘部署设计的Sprinklr-Edge-OCR系统,并对五种主流LVLMs(InternVL、Qwen、GOT OCR、LLaMA、MiniCPM)和两种传统OCR系统(Sprinklr-Edge-OCR、SuryaOCR)在包含54种语言的自建双人工标注数据集上进行大规模对比评估。评测涵盖准确率、语义一致性、语言覆盖、计算效率(延迟、内存、GPU占用)及部署成本。通过边缘场景分析,评估了仅在CPU环境下的表现。结果显示,Qwen精度最高(0.54),但Sprinklr-Edge-OCR整体F1得分最高(0.46),平均处理速度达0.17秒/图,比LVLM快35倍,成本低至0.006美元/千图,不足其1%。研究证明,在边缘部署中,传统系统因低算力、低延迟和极高性价比仍具优势。

原文摘要 · Abstract (English)

Optical Character Recognition (OCR) in multilingual, noisy, and diverse real-world images remains a significant challenge for optical character recognition systems. With the rise of Large Vision-Language Models (LVLMs), there is growing interest in their ability to generalize and reason beyond fixed OCR pipelines. In this work, we introduce Sprinklr-Edge-OCR, a novel OCR system built specifically optimized for edge deployment in resource-constrained environments. We present a large-scale comparative evaluation of five state-of-the-art LVLMs (InternVL, Qwen, GOT OCR, LLaMA, MiniCPM) and two traditional OCR systems (Sprinklr-Edge-OCR, SuryaOCR) on a proprietary, doubly hand annotated dataset of multilingual (54 languages) images. Our benchmark covers a broad range of metrics including accuracy, semantic consistency, language coverage, computational efficiency (latency, memory, GPU usage), and deployment cost. To better reflect real-world applicability, we also conducted edge case deployment analysis, evaluating model performance on CPU only environments. Among the results, Qwen achieved the highest precision (0.54), while Sprinklr-Edge-OCR delivered the best overall F1 score (0.46) and outperformed others in efficiency, processing images 35 faster (0.17 seconds per image on average) and at less than 0.01 of the cost (0.006 USD per 1,000 images) compared to LVLM. Our findings demonstrate that the most optimal OCR systems for edge deployment are the traditional ones even in the era of LLMs due to their low compute requirements, low latency, and very high affordability.

OCR边缘计算多语言大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。