arXiv:2505.23008cs.CVcs.AI2025-05被引 2

构建匈牙利语文档问答数据集,提升低资源语言视觉问答性能

Synthetic Document Question Answering in Hungarian

  • 基于匈牙利网页文档生成大规模合成数据集
  • 微调后模型在匈牙利语文档问答上准确率提升7.2%
  • 适合多语言文档理解与低资源语言研究者使用

现代视觉语言模型在英文文档视觉问答任务中已接近饱和精度,但在低资源语言中仍具挑战性,主要因缺乏合适的训练与评估数据。本文聚焦匈牙利语(互联网资源约第17位),提出可扩展的数据集构建方法。我们构建了两个文档视觉问答数据集:HuDocVQA(大规模合成数据)和HuDocVQA-manual(小规模人工标注数据),均源自Common Crawl的匈牙利文档。通过多轮质量过滤与去重,使HuDocVQA达到人类水平。此外,我们发布包含11.7万页匈牙利语PDF及其转录文本的HuCCPDF数据集,可用于训练匈牙利语OCR模型。实验表明,将这些数据混合微调Llama 3.2 11B Instruct模型,在HuDocVQA上的准确率提升7.2%。所有数据与代码将公开,以推动多语言文档问答研究。

原文摘要 · Abstract (English)

Modern VLMs have achieved near-saturation accuracy in English document visual question-answering (VQA). However, this task remains challenging in lower resource languages due to a dearth of suitable training and evaluation data. In this paper we present scalable methods for curating such datasets by focusing on Hungarian, approximately the 17th highest resource language on the internet. Specifically, we present HuDocVQA and HuDocVQA-manual, document VQA datasets that modern VLMs significantly underperform on compared to English DocVQA. HuDocVQA-manual is a small manually curated dataset based on Hungarian documents from Common Crawl, while HuDocVQA is a larger synthetically generated VQA data set from the same source. We apply multiple rounds of quality filtering and deduplication to HuDocVQA in order to match human-level quality in this dataset. We also present HuCCPDF, a dataset of 117k pages from Hungarian Common Crawl PDFs along with their transcriptions, which can be used for training a model for Hungarian OCR. To validate the quality of our datasets, we show how finetuning on a mixture of these datasets can improve accuracy on HuDocVQA for Llama 3.2 11B Instruct by +7.2%. Our datasets and code will be released to the public to foster further research in multilingual DocVQA.

文档问答多语言合成数据匈牙利语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。