用多模态大模型实现零样本多页手写文档转录,提升跨页上下文利用效率。
Judge a Book by its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription
- 设计跨页共享内容的提示策略,融合OCR与多模态大模型优势
- 在新构建的Malvern-Hills数据集上,性能超越现有方法12.3%
- 适合无标注数据场景下的历史文献、档案等多页手写文本处理
手写文字识别(HTR)仍是难题。现有方法依赖标注数据微调,难以获取;或采用零样本工具如OCR引擎和多模态大模型(MLLM)。MLLM在端到端转录和后处理中表现良好,但对多页文档的跨页上下文利用不足。多数手写文档为多页,页面间共享语义和书写风格,但现有方法通常逐页处理,忽略上下文。本文研究结合OCR、LLM后处理与MLLM端到端转录的多种方法,针对零样本多页手写文档转录任务。我们基于现有单页数据集构建新基准,包括新数据集Malvern-Hills。提出OCR+PAGE-1与OCR+PAGE-N两种提示策略,在跨页共享内容的同时降低提示复杂度,性能优于现有方法。
原文摘要 · Abstract (English)
Handwriting text recognition (HTR) remains a challenging task. Existing approaches require fine-tuning on labeled data, which is impractical to obtain for real-world problems, or rely on zero-shot tools such as OCR engines and multi-modal LLMs (MLLMs). MLLMs have shown promise both as end-to-end transcribers and as OCR post-processors, but to date there is little empirical research evaluating different MLLM prompting strategies for HTR, particularly for the case of multi-page documents. Most handwritten documents are multi-page, and share context such as semantic content and handwriting style across pages, yet MLLMs are typically used for transcription at the page level, meaning they throw away this shared context. They are also typically used as either text-only post-processors or image-only OCR alternatives, rather than leveraging multiple modes. This paper investigates a suite of methods combining OCR, LLM post-processing and MLLM end-to-end transcription, for the task of zero-shot multi-page handwritten document transcription. We introduce a benchmark for this task from existing single-page datasets, including a new dataset, Malvern-Hills. Finally, we introduce OCR+PAGE-1 and OCR+PAGE-N, prompting strategies for multi-page transcription that outperform existing methods by sharing content across pages while minimizing prompt complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。