用邻居启发工具链提升视觉大模型识别低质表格的能力
Enhancing Table Recognition with Vision LLMs: A Benchmark and Neighbor-Guided Toolchain Reasoner
- 借鉴邻近样本工具选择经验,动态调用轻量工具优化图像输入
- 在公开数据集上显著提升原始视觉大模型的表格识别准确率
- 适合需要高精度表格解析的科研与工业场景
预训练基础模型在表格理解与推理任务中取得显著进展,但利用视觉大语言模型(VLLMs)识别非结构化表格的结构与内容仍研究不足。为弥补这一差距,我们基于分层设计思想构建了一个无需训练的评估基准,用于衡量VLLMs的识别能力。深入评估发现,低质量图像输入是识别过程中的主要瓶颈。受此启发,我们提出邻居引导工具链推理框架(NGTR),通过集成多种轻量级视觉操作工具来缓解低质图像问题。具体而言,将相似邻近样本的工具选择经验迁移至当前输入,并设计反思模块监督工具调用过程。在多个公开数据集上的大量实验表明,该方法显著提升了原始VLLMs的识别性能。我们认为该基准与框架可为表格识别提供新解决方案。
原文摘要 · Abstract (English)
Pre-trained foundation models have recently made significant progress in table-related tasks such as table understanding and reasoning. However, recognizing the structure and content of unstructured tables using Vision Large Language Models (VLLMs) remains under-explored. To bridge this gap, we propose a benchmark based on a hierarchical design philosophy to evaluate the recognition capabilities of VLLMs in training-free scenarios. Through in-depth evaluations, we find that low-quality image input is a significant bottleneck in the recognition process. Drawing inspiration from this, we propose the Neighbor-Guided Toolchain Reasoner (NGTR) framework, which is characterized by integrating diverse lightweight tools for visual operations aimed at mitigating issues with low-quality images. Specifically, we transfer a tool selection experience from a similar neighbor to the input and design a reflection module to supervise the tool invocation process. Extensive experiments on public datasets demonstrate that our approach significantly enhances the recognition capabilities of the vanilla VLLMs. We believe that the benchmark and framework could provide an alternative solution to table recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。