arXiv:2606.02162cs.CVcs.AI2026-06

对比四种模型,发现图像信息对文档分类最关键

Multimodal Approaches for Visually-Rich Document Type Classification: A Comparative Analysis

论文配图:Multimodal Approaches for Visually-Rich Document Type Classification: A Comparative Analysis
图 1 · 摘自论文原文
  • 统一框架下比较Transformer与大模型的多模态设计
  • 图像特征贡献最大,光学字符识别文本为辅助
  • 适合需要布局分析的文档分类任务研究者

视觉丰富的文档类型分类仍具挑战性,因关键信息分散于文本、图像和版式模态中。现有方法采用多样化的多模态建模策略,导致架构异质性强,难以系统比较。现有对比研究也常使用异构评估设置,进一步阻碍了进展评估。为此,本文在统一实验框架下,系统分析基于Transformer与大语言模型的多模态设计策略。具体评估了四种代表性模型(LayoutLMv3、Donut、Qwen3-VL-32B-Instruct 和 Qwen3-32B)在 RVL-CDIP 基准上的表现,重点对比依赖与不依赖 OCR 的方法。结果表明:专用多模态 Transformer 在视觉丰富且版式复杂的文档上优于大模型;图像信息对可靠分类贡献最强,而 OCR 提取的文本仅起辅助作用。研究揭示了多模态处理对具有明显版式结构文档的重要性,为模型选择与特征组合提供了实证依据。

原文摘要 · Abstract (English)

Document type classification in visually rich documents remains challenging, as relevant information is distributed across textual, visual, and layout modalities. To capture this complexity, current approaches rely on diverse multimodal modeling strategies, resulting in heterogeneous architectures that complicate systematic comparison. This variability is also reflected in existing comparative studies, which often rely on heterogeneous evaluation setups, further complicating systematic comparison and making it difficult to assess progress. To address these limitations, this work provides a structured analysis of multimodal design strategies across transformer- and LLM-based architectures, combined with a controlled empirical comparison within a unified experimental framework. Specifically, four representative models (LayoutLMv3, Donut, Qwen3-VL-32B-Instruct, and Qwen3-32B) are evaluated on the RVL-CDIP benchmark to systematically analyze the contributions of text, image, and layout information for document type classification, with a particular focus on contrasting OCR-dependent and OCR-free approaches. The results show that specialized multimodal Transformers outperform LLM-based approaches on visually rich and layout-intensive documents. Image information contributes most strongly to reliable classification, while OCR-derived text provides useful but secondary support. These findings highlight that multimodal processing remains essential for documents with pronounced layout structure. Overall, the study provides a systematic basis for comparing multimodal architectures and offers practical guidance for selecting effective feature combinations and model designs for document type classification.

文档分类多模态版式分析视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。