arXiv:2505.17625cs.CLcs.CV2025-05中稿 · IIAI AAI 2025, the…被引 4

用布局和文本增强视觉语言模型,提升日文财报表格问答准确率

Enhancing Large Vision-Language Models with Layout Modality for Table Question Answering on Japanese Annual Securities Reports

  • 融合表格内文字与版式特征,改进视觉语言模型理解力
  • 在日文财报表格问答任务中显著提升性能,无需结构化输入
  • 适合金融文档智能解析、多模态模型研究者使用

随着大语言模型(LLMs)的发展和检索增强生成(RAG)的兴起,理解表格结构变得愈发重要。尤其在证券报告等金融领域,对表格内容的高精度问答需求迫切。然而,表格格式多样(如HTML、图像、纯文本),难以保留和提取结构信息。因此,多模态大模型对鲁棒且通用的表格理解至关重要。尽管大型视觉语言模型(LVLMs)具有潜力,但其在理解文档中字符及其空间关系方面仍存在挑战。本文提出一种方法,通过引入表格内的文本内容和布局特征,增强基于LVLM的表格理解能力。实验表明,这些辅助模态显著提升了性能,使模型能在不依赖显式结构化输入的前提下,稳健解析复杂文档布局。

原文摘要 · Abstract (English)

With recent advancements in Large Language Models (LLMs) and growing interest in retrieval-augmented generation (RAG), the ability to understand table structures has become increasingly important. This is especially critical in financial domains such as securities reports, where highly accurate question answering (QA) over tables is required. However, tables exist in various formats-including HTML, images, and plain text-making it difficult to preserve and extract structural information. Therefore, multimodal LLMs are essential for robust and general-purpose table understanding. Despite their promise, current Large Vision-Language Models (LVLMs), which are major representatives of multimodal LLMs, still face challenges in accurately understanding characters and their spatial relationships within documents. In this study, we propose a method to enhance LVLM-based table understanding by incorporating in-table textual content and layout features. Experimental results demonstrate that these auxiliary modalities significantly improve performance, enabling robust interpretation of complex document layouts without relying on explicitly structured input formats.

表格问答视觉语言模型金融文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。