arXiv:2502.19202cs.CL2025-02中稿 · IJDAR被引 3

首个越南语收据文档VQA数据集,融合版面信息提升问答准确率。

LiGT: Layout-infused Generative Transformer for Visual Question Answering on Vietnamese Receipts

  • 设计布局感知的生成式Transformer,用语言模型嵌入处理版面信息。
  • 在9000+张收据上实现6万+问答对,性能媲美顶尖基线模型。
  • 强调多模态融合必要性,生成式架构优于纯编码器模型。

文档视觉问答(Document VQA)要求多模态系统综合处理文本、版面和视觉信息以给出恰当答案。近年来随着文档数量激增和数字化需求上升,该任务日益重要。然而,现有大多数文档VQA数据集集中在英语等高资源语言。本文提出首个面向收据的越南语大规模文档VQA数据集ReceiptVQA,包含9,000+张收据图像和60,000+组人工标注的问答对。同时,我们提出LiGT(Layout-infused Generative Transformer)——一种布局感知的编码器-解码器架构,通过语言模型嵌入层处理版面信息,减少额外神经模块使用。在ReceiptVQA上的实验表明,该架构表现优异,达到与先进基线相当的性能。分析发现:生成式架构显著优于仅编码器模型;多模态融合对本任务至关重要,尽管语言模型的语义理解能力已很关键。我们希望本工作能推动越南语文档VQA发展,促进越南语多模态研究生态建设。

原文摘要 · Abstract (English)

Document Visual Question Answering (Document VQA) challenges multimodal systems to holistically handle textual, layout, and visual modalities to provide appropriate answers. Document VQA has gained popularity in recent years due to the increasing amount of documents and the high demand for digitization. Nonetheless, most of document VQA datasets are developed in high-resource languages such as English. In this paper, we present ReceiptVQA (\textbf{Receipt} \textbf{V}isual \textbf{Q}uestion \textbf{A}nswering), the initial large-scale document VQA dataset in Vietnamese dedicated to receipts, a document kind with high commercial potentials. The dataset encompasses \textbf{9,000+} receipt images and \textbf{60,000+} manually annotated question-answer pairs. In addition to our study, we introduce LiGT (\textbf{L}ayout-\textbf{i}nfused \textbf{G}enerative \textbf{T}ransformer), a layout-aware encoder-decoder architecture designed to leverage embedding layers of language models to operate layout embeddings, minimizing the use of additional neural modules. Experiments on ReceiptVQA show that our architecture yielded promising performance, achieving competitive results compared with outstanding baselines. Furthermore, throughout analyzing experimental results, we found evident patterns that employing encoder-only model architectures has considerable disadvantages in comparison to architectures that can generate answers. We also observed that it is necessary to combine multiple modalities to tackle our dataset, despite the critical role of semantic understanding from language models. We hope that our work will encourage and facilitate future development in Vietnamese document VQA, contributing to a diverse multimodal research community in the Vietnamese language.

文档VQA越南语多模态收据识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。