arXiv:2507.15655cs.CV2025-07被引 1

构建首个多语言手写文档VQA基准,推动跨语言手写识别发展

HW-MLVQA: Elucidating Multilingual Handwritten Document Understanding with a Comprehensive VQA Benchmark

  • 构建包含1600页手写文档的多语言VQA数据集
  • 提供文本、图像及图文融合三种模态评估框架
  • 适合手写识别、多模态模型研究者使用

多语言视觉问答(MLVQA)基准的兴起提升了大语言模型(LLMs)与多模态大模型的能力,使其能更好地捕捉不同语言中的语言细微差异与视觉复杂性。然而,现有模型在处理多样化手写文档时仍表现不足。本文提出HW-MLVQA,一个专为多语言手写文档理解设计的前沿视觉问答基准。该基准包含1,600页真实手写文档,配套2,400组问题与答案,并提供覆盖文本、图像及图文融合三种模态的评估框架。为模拟无真实文本转录的真实场景,基准支持对自研与开源OCR模型的严格评测。该工作旨在推动多语言手写文档理解的关键进展,促进该领域的创新与学术研究。

原文摘要 · Abstract (English)

The proliferation of MultiLingual Visual Question Answering (MLVQA) benchmarks augments the capabilities of large language models (LLMs) and multi-modal LLMs, thereby enabling them to adeptly capture the intricate linguistic subtleties and visual complexities inherent across diverse languages. Despite its potential, the current MLVQA model struggles to fully utilize its capabilities when dealing with the extensive variety of handwritten documents. This article delineates HW-MLVQA, an avant-garde VQA benchmark meticulously crafted to mitigate the dearth of authentic Multilingual Handwritten document comprehension. HW-MLVQA encompasses an extensive collection of 1,600 handwritten Pages complemented by 2,400 question-answers. Furthermore, it provides a robust benchmark evaluation framework spanning three distinct modalities: text, image, and an integrated image & text modality. To simulate authentic real-world contexts devoid of ground truth textual transcriptions, we facilitates a rigorous assessment of proprietary and open-source OCR models. The benchmark aspires to facilitate pivotal advancements in multilingual handwritten document interpretation, fostering innovation and scholarly inquiry within this specialized domain.

手写识别多语言VQA多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。