arXiv:2503.23667cs.CV2025-03被引 4

测试多模态大模型在不同图像分辨率下的单字符识别能力

Context-Independent OCR with Multimodal LLMs: Effects of Image Resolution and Visual Complexity

  • 用单字符图像测试大模型脱离上下文的识别效果
  • 300 ppi时表现媲美传统OCR,低于150 ppi时性能骤降
  • 视觉复杂度对识别错误影响极小,适合高精度场景

由于在图像描述、文档分析和自动化内容生成等任务中的高通用性,多模态大语言模型(Multimodal LLMs)在多个工业领域受到广泛关注。尤其在光学字符识别(OCR)任务中,其表现已超越专用模型。然而,其在不同图像条件下的性能仍缺乏充分研究,且依赖上下文线索可能导致单个字符识别不准确。本文通过使用具有不同视觉复杂度的单字符图像,考察上下文无关的OCR任务,以确定准确识别的条件。结果表明,多模态LLMs在约300 ppi时可达到与传统OCR方法相当的性能,但在低于150 ppi时显著下降。此外,视觉复杂度与误识别之间相关性极弱,而专用OCR模型则无此相关性。这些发现表明,图像分辨率和视觉复杂度在需要精确字符级准确性的多模态LLM OCR应用中至关重要。

原文摘要 · Abstract (English)

Due to their high versatility in tasks such as image captioning, document analysis, and automated content generation, multimodal Large Language Models (LLMs) have attracted significant attention across various industrial fields. In particular, they have been shown to surpass specialized models in Optical Character Recognition (OCR). Nevertheless, their performance under different image conditions remains insufficiently investigated, and individual character recognition is not guaranteed due to their reliance on contextual cues. In this work, we examine a context-independent OCR task using single-character images with diverse visual complexities to determine the conditions for accurate recognition. Our findings reveal that multimodal LLMs can match conventional OCR methods at about 300 ppi, yet their performance deteriorates significantly below 150 ppi. Additionally, we observe a very weak correlation between visual complexity and misrecognitions, whereas a conventional OCR-specific model exhibits no correlation. These results suggest that image resolution and visual complexity may play an important role in the reliable application of multimodal LLMs to OCR tasks that require precise character-level accuracy.

OCR多模态大模型图像质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。