arXiv:2603.22815cs.CVcs.AI2026-03被引 1

聚焦关键区域,提升图文理解效率与准确率

Focus, Don't Prune: Identifying Instruction-Relevant Regions for Information-Rich Image Understanding

  • 先定位指令相关图像区域,再提取细粒度特征
  • 在多页文档等复杂图像上准确率显著提升
  • 适合需要高效处理信息密集图像的场景

大型视觉语言模型(LVLM)在多模态任务中表现优异,但处理信息密集图像(如信息图、文档布局)时需生成大量视觉标记,导致计算开销大。为此,我们提出PinPoint框架,分两阶段进行:首先基于视觉输入和文本指令定位指令相关区域,再精细提取特征以提升推理能力与效率。核心是“指令-区域对齐”机制,并构建了针对InfographicVQA、MultiPageDocVQA、SinglePageDocVQA等挑战性VQA基准的新标注数据集,提供更丰富的监督信号。实验表明,PinPoint不仅在准确率上优于现有方法,还通过减少无关视觉标记显著降低计算开销。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) have shown strong performance across various multimodal tasks by leveraging the reasoning capabilities of Large Language Models (LLMs). However, processing visually complex and information-rich images, such as infographics or document layouts, requires these models to generate a large number of visual tokens, leading to significant computational overhead. To address this, we propose PinPoint, a novel two-stage framework that first identifies instruction-relevant image regions and then refines them to extract fine-grained visual features for improved reasoning and efficiency. Central to our approach is the Instruction-Region Alignment, which localizes relevant regions using both visual input and textual instructions. We further introduce new annotations that provide richer ground-truth supervision for instruction-relevant regions across challenging VQA benchmarks: InfographicVQA, MultiPageDocVQA, and SinglePageDocVQA. Experimental results show that PinPoint not only achieves superior accuracy compared to existing methods but also reduces computational overhead by minimizing irrelevant visual tokens.

视觉语言模型图像理解高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。