arXiv:2604.11970cs.CVcs.AI2026-04ACL

构建印尼文档跨语言表格理解基准,推动多语言文档智能解析研究。

INDOTABVQA: A Benchmark for Cross-Lingual Table Understanding in Bahasa Indonesia Documents

论文配图:INDOTABVQA: A Benchmark for Cross-Lingual Table Understanding in Bahasa Indonesia Documents
图 1 · 摘自论文原文
  • 构建含1593张真实文档图像的跨语言表格VQA数据集,覆盖多种排版风格。
  • 在复杂表格和低资源语言上,现有模型准确率差距显著,最高提升17.8%。
  • 添加表格区域坐标可提高4-7%性能,适合研究多语言文档理解的学者。

我们提出INDOTABVQA,一个面向印尼语文档图像的跨语言表格视觉问答评估基准。数据集包含1593张真实文档图像,涵盖有边框、无边框和彩色三种视觉样式,每张图像含一个或多个表格,配套1593组问答对,覆盖印尼语、英语、印地语和阿拉伯语四种语言。该数据集支持在单语(印尼文档+印尼问题)与跨语言(印尼文档+其他语言问题)场景下评估视觉语言模型(VLMs)。我们测试了Qwen2.5-VL、Gemma-3、LLaMA-3.2等主流开源VLM及GPT-4o,发现其在结构复杂表格和低资源语言上存在显著性能差距。在本数据集上微调3B小型模型和LoRA微调7B模型,准确率分别提升11.6%和17.8%。额外提供表格区域坐标作为输入,性能进一步提升4-7%,表明空间先验对表格推理至关重要。研究强调语言多样性和领域特异性数据的重要性,并证明针对性微调能显著提升模型在专业文档理解任务上的表现。完整数据集可在Hugging Face获取:https://huggingface.co/datasets/NusaBharat/INDOTABVQA

原文摘要 · Abstract (English)

We introduce INDOTABVQA, a benchmark for evaluating cross-lingual Table Visual Question Answering (VQA) on real-world document images in Bahasa Indonesia. The dataset comprises 1,593 document images across three visual styles (bordered, borderless, and colorful) with one or more than one tables, and 1,593 question-answer sets in four languages: Bahasa Indonesia, English, Hindi, and Arabic. This enables evaluation of Vision-Language Models (VLMs) in both monolingual (Bahasa documents with Bahasa questions) and cross-lingual settings (Bahasa documents with questions in other languages). We benchmark leading open-source VLMs (Qwen2.5-VL, Gemma-3, LLaMA-3.2) and GPT-4o and reveal substantial performance gaps, particularly on structurally complex tables and in low-resource languages. Fine-tuning a compact 3B and LoRA-finetuned 7B model on our dataset yields 11.6% and 17.8% improvements in accuracy. Providing explicit table region coordinates as additional input further improves performance by 4-7%, demonstrating the value of Spatial priors for table-based reasoning. Our findings underscore the importance of language-diverse, domain-specific datasets and demonstrate that targeted fine-tuning can significantly enhance VLM performance on specialized document understanding tasks. INDOTABVQA provides a valuable resource for advancing research in cross-lingual, structure-aware document understanding, especially in underrepresented regions of the world. Full dataset can be accessed in huggingface at: https://huggingface.co/datasets/NusaBharat/INDOTABVQA}

表格理解跨语言文档解析多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。