arXiv:2608.22959cs.CVcs.AI2026-08

挑战手写文本理解的基准,揭示大模型依赖语言先验的弱点

WildHandBench: A Benchmark for Handwritten Text Understanding that Challenges MLLMs and Humans

论文配图:WildHandBench: A Benchmark for Handwritten Text Understanding that Challenges MLLMs and Humans
图 1 · 摘自论文原文
  • 构建涵盖多结构、多语言、真实场景的手写文档基准集
  • 最佳模型仅达71.85%准确率,人类为77.09%且差距小
  • 提出先验驱动误差度量,揭示模型对语言先验的系统性依赖

尽管当前最佳模型在OmniDocBench上的印刷文档解析整体准确率达96.34%,但现有模型处理复杂手写文档的能力仍缺乏充分评估。现有基准多关注孤立文本或公式,忽视手写表格与真实世界退化,且仅报告综合准确率而不分析失败原因。我们提出WildHandBench,包含500份手写文档,覆盖三种结构(自由文本、表格、公式)、四种语言及九种真实场景。引入先验驱动误差(PDE)度量,量化错误是否源于语言先验而非视觉证据。评估18个先进模型及校准的人类基线发现:(1) 最佳模型整体准确率为71.85%;(2) 人类表现优于所有模型(77.09% vs. 71.85%),但差距较小;(3) 模型错误中63-91%为先验驱动,而人类仅为49%,暴露了模型对语言先验的系统性依赖,传统准确率指标无法捕捉此问题。

原文摘要 · Abstract (English)

While the top model on OmniDocBench now reaches 96.34% overall on printed-document parsing, the ability of current models to handle challenging handwritten documents remains largely uncharacterized. Existing benchmarks focus on isolated text or formulas, overlook handwritten tables and real-world degradation, and report aggregate accuracy without explaining why models fail. We present WildHandBench, a benchmark containing 500 handwritten documents across three structures (free text, tables, formulas), four languages, and nine real-world scenarios. We introduce a Prior-Driven Error (PDE) metric that quantifies whether errors originate from language priors rather than visual evidence. Evaluating 18 state-of-the-art models together with calibrated human baselines, we find: (1) the best model achieves only 71.85% overall; (2) humans outperform all models yet the gap is narrow (77.09% vs. 71.85%); and (3) model errors are qualitatively different from human errors -- 63-91% of model errors are prior-driven versus only 49% for humans, exposing systematic reliance on language priors that conventional accuracy metrics cannot capture.

手写识别多模态模型基准测试语言先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。