评测大模型在文字定位与推理上的能力,发现多数模型表现不佳。
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
- 构建包含31种场景的多任务中文英文基准测试
- 覆盖1万组人工验证题,1500张手动标注测试图
- 揭示模型在手写体、布局理解等五类问题上普遍薄弱
评估大型多模态模型(LMMs)的光学字符识别(OCR)能力日益受到关注。现有基准虽展示出LMM在文本识别上的优异表现,但在文本定位、手写内容提取及逻辑推理等挑战性任务上仍研究不足。为此,我们推出OCR-Bench v2,一个大规模双语文本为中心的基准测试,涵盖当前最全面的任务集(较前代多4倍),覆盖31种多样场景,并提供详尽的评估指标。该数据集包含10,000组经人工验证的问答对,且难样本比例高。此外,我们构建了含1,500张人工标注图像的私有测试集。公共与私有测试集间一致的评估趋势验证了OCR-Bench v2的可靠性。对主流LMM的系统评测发现,多数模型得分低于50(满分100),存在五类缺陷:不常见文本识别、细粒度感知、版面感知、复杂元素解析及逻辑推理能力不足。
原文摘要 · Abstract (English)
Scoring the Optical Character Recognition (OCR) capabilities of Large Multimodal Models (LMMs) has witnessed growing interest. Existing benchmarks have highlighted the impressive performance of LMMs in text recognition; however, their abilities in certain challenging tasks, such as text localization, handwritten content extraction, and logical reasoning, remain underexplored. To bridge this gap, we introduce OCRBench v2, a large-scale bilingual text-centric benchmark with currently the most comprehensive set of tasks (4x more tasks than the previous multi-scene benchmark OCRBench), the widest coverage of scenarios (31 diverse scenarios), and thorough evaluation metrics, with 10,000 human-verified question-answering pairs and a high proportion of difficult samples. Moreover, we construct a private test set with 1,500 manually annotated images. The consistent evaluation trends observed across both public and private test sets validate the OCRBench v2's reliability. After carefully benchmarking state-of-the-art LMMs, we find that most LMMs score below 50 (100 in total) and suffer from five-type limitations, including less frequently encountered text recognition, fine-grained perception, layout perception, complex element parsing, and logical reasoning. The project website is at: https://99franklin.github.io/ocrbench_v2/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。