arXiv:2604.16504cs.CVcs.LG2026-04

评测17个大模型对复杂手写表单的数字化能力,顶尖模型准确率达85%。

From Handwriting to Structured Data: Benchmarking AI Digitisation of Handwritten Forms

论文配图:From Handwriting to Structured Data: Benchmarking AI Digitisation of Handwritten Forms
图 1 · 摘自论文原文
  • 用真实医疗表单测试多模态大模型,含手写、打印、日期等混合内容
  • GPT-5.4在日期提取上最准,幻觉率仅6%;Gemini 3.1整体表现最佳
  • 优化提示词可大幅提升精确率,但对加权指标影响小

手工数字化结构化手写文档耗时且成本高。本文针对一个包含日期、印刷文本、手写回答及显著变异性的真实医疗表单,对17个前沿多模态大模型和开源模型进行基准测试。较小或较旧的模型表现不佳,而最新的谷歌与OpenAI模型在离散字段上的准确率约为85%,加权F1得分约90%,尽管响应极具挑战性。任务特定优势明显:GPT-5.4在噪声日期提取中表现最佳,幻觉率最低(6%);Claude Sonnet 4.6在格式化字段(日期与数值)上平均表现最好;Gemini 3.1整体最优,自由文本错误率最低(WER=0.50,CER=0.31),离散分类指标最强。进一步表明,提示词优化使宏精度、召回率和F1提升超60%,但对加权指标改善有限(仅约2–5%)。结果表明,多模态大模型的快速进步为复杂手写流程的全自动数字化提供了可行路径,尤其适用于低收入和中等收入国家。

原文摘要 · Abstract (English)

Manual digitisation of structured handwritten documents is slow and costly. We benchmark 17 leading frontier multi-modal large language models and open-source models against a very challenging real-world medical form that mixes dates; structured, printed text; hand-written responses and significant variability challenges. None of the smaller or older models perform well but the latest Google and OpenAI models reach accuracies around $85\%$ with weighted F1 scores $\simeq 90\%$ across the discrete or predefined fields despite the very challenging nature of the responses. Clear task specific strengths emerge: GPT 5.4 excels in noisy date extraction as well as reliability with the lowest hallucination rate ($6\%$). Claude Sonnet 4.6 had the best average performance across formatted fields (dates and numerical values), while Gemini 3.1 delivered the best overall performance, with the lowest free text error rates (WER = $0.50$ and CER = $0.31$) and the strongest results across discrete classification metrics. We further show that prompt optimisation dramatically improves macro precision, recall and F1 by over $60\%$, but has little impact on weighted metrics (only $\sim2-5\%$ improvement). These results provide evidence that the rapid improvements of multimodal large language models offer a compelling pathway toward fully automated digitisation of complex handwritten workflows that is particularly relevant in low- and middle-income countries.

手写识别多模态医疗数字化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。