对比8个视觉语言模型在三类文档上的表现,发现小模型微调更省力,视觉理解是关键瓶颈。
Comparative Study of Domain-adapted VLMs for General Document Visual Question Answering

- 系统评估8个开源VLM在工业文档、信息图、幻灯片上的表现
- 小模型经微调后性能提升显著,50样本即可快速适应新领域
- 视觉复杂布局下模型表现下降,说明视觉理解是主要瓶颈
文档视觉问答(DocVQA)是一项复杂的多模态挑战,要求模型融合文档的视觉、文本和版面信息。尽管视觉语言模型(VLMs)在图文任务中表现出色,但其在不同文档领域的鲁棒性和可迁移性仍待深入探索。本研究对8个开源预训练VLMs在三类文档域——不同类型工业文档、信息图、演示文稿——上的表现进行了全面评估。通过零样本、全监督微调(跨/同数据集)、少样本知识迁移等实验,发现大型VLM在结构化布局上具备强零样本能力,但在视觉复杂的图表与幻灯片中表现明显下降。虽然参数量是性能的关键因素,但小模型在监督微调中获得更高的相对收益。跨域与少样本实验表明,视觉理解是制约性能的主要瓶颈,而非模型知识不足。仅用50个目标域样本进行微调,模型即可快速适应目标文档,甚至在某些情况下超越全监督模型。
原文摘要 · Abstract (English)
Document Visual Question Answering (DocVQA) presents a complex multimodal challenge, requiring models to exploit visual, textual, and layout information from documents. Although Vision-Language Models (VLMs) have shown remarkable performance in text-vision tasks, their robustness and transferability to different document domains remains underexplored. In this study, we present a comprehensive evaluation of 8 open-source pretrained VLMs on DocVQA in three different document domains: industrial documents of varying type, infographics, and presentation slides. We systematically assess model performance under zero-shot evaluations, fully supervised finetuning with inter- and intra-dataset evaluations, and few-shot learning evaluations of knowledge transfer between domains. Our findings demonstrate that while large pretrained VLMs possess strong zero-shot baselines for structured layouts, their performance strongly decreases on visually complex layouts of infographics and slides. Although parameter scaling is a dominant factor on performance, supervised finetuning yields higher relative gains in smaller architectures. Furthermore, our cross-domain and few-shot experiments show that visual understanding is the main bottleneck for DocVQA, not a lack of knowledge from the VLMs. Using 50 target domain samples, the models finetuned in DocVQA with datasets of different domains rapidly adapt to the target domain documents, even surpassing their fully supervised counterparts in some cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。