构建手写文字与公式识别诊断基准,评测大模型真实手写场景表现。
OmniHandwritingOCR: A Diagnostic Benchmark for Evaluating Multimodal LLMs in Handwritten OCR Scenarios

- 设计多语言、复杂公式等12个子集,覆盖真实手写多样场景
- 13个模型在复杂多行公式上性能骤降,平均准确率不足60%
- 发现生成模型常凭空编造视觉无依据的纠错内容
多模态大语言模型(MLLM)日益用于文档与知识处理中的光学字符识别(OCR),但其对真实手写文本的识别能力仍缺乏深入评估。现有基准多聚焦印刷体或单行清晰输入,难以覆盖多语言手写、书写错误及结构复杂的数学表达等真实场景。本文提出OmniHandwritingOCR,一个用于评估MLLM与OCR系统在手写场景下的诊断性基准。该基准涵盖六项子任务与十二个数据子集,共包含77.57万张标注图像,来自公开数据集和新采集的学生手写样本。核心部分为分难度层级的多行公式语料库,用于测试模型在结构复杂度递增下的鲁棒性。我们采用统一协议与五种互补指标,评估了十三个开源与闭源系统。结果表明,当前系统距离忠实转录仍有显著差距:在复杂多行公式上性能急剧下降,模型排名随语言与公式类型变化,且多个生成模型会虚构出视觉上无支持的合理修正。OmniHandwritingOCR为诊断多模态模型在手写识别中出现的语言、内容、结构与视觉对齐失败提供了挑战性测试平台。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) are increasingly used as OCR systems in document and knowledge-processing pipelines, but their ability to faithfully read real handwriting remains underexplored. Existing OCR benchmarks focus largely on printed text or clean single-line inputs, leaving limited coverage of realistic handwritten OCR scenarios such as multilingual handwriting, writer errors, and structurally complex mathematical expressions. We introduce OmniHandwritingOCR, a diagnostic benchmark for evaluating MLLMs and OCR systems on handwritten OCR. It covers handwritten text recognition and handwritten mathematical expression recognition across six subtasks and twelve subsets, totaling 77.57K labeled images from public datasets and newly collected student writings. A key component is a difficulty-stratified multi-line formula corpus designed to test robustness under increasing structural complexity. We evaluate thirteen open- and closed-source systems with five complementary metrics under a unified protocol. Results show that current systems remain far from faithful transcription: performance drops sharply on complex multi-line formulas, model rankings vary across language and formula settings, and several generative models hallucinate plausible but visually unsupported corrections. OmniHandwritingOCR provides a challenging testbed for diagnosing language, content, structural, and visual-grounding failure modes of multimodal models in handwritten OCR scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。