arXiv:2508.05053cs.CV2025-08ACL

测试大模型找文档细节能力,提出高效聚焦方法

Finding Needles in Images: Can Multimodal LLMs Locate Fine Details?

  • 通过智能选区与高斯注意力模拟人类聚焦搜索
  • 在复杂版式中定位细节准确率提升显著
  • 适合需要精准提取信息的文档分析场景

尽管多模态大语言模型在文档理解任务中表现优异,但其在复杂文档中定位和推理细粒度信息的能力仍缺乏研究。例如,在餐厅菜单中查找特定营养信息,或在长篇报纸文章中识别免责声明——这些任务要求对细微但关键的信息保持高度关注,类似于在图像中寻找针头(Finding Needles in Images, NiM)。为此,我们构建了NiM基准,涵盖新闻、菜单、讲义等多种真实文档,专门评估模型在复杂任务中的表现。在此基础上,我们提出Spot-IT方法,通过智能补丁选择与高斯注意力机制,模仿人类在文档中缩放聚焦的视觉行为。大量实验揭示了当前多模态大模型在细粒度文档理解中的能力和局限性,同时验证了所提方法的有效性。Spot-IT在需要从复杂布局中精确提取细节的任务中显著优于基线方法。

原文摘要 · Abstract (English)

While Multi-modal Large Language Models (MLLMs) have shown impressive capabilities in document understanding tasks, their ability to locate and reason about fine-grained details within complex documents remains understudied. Consider searching a restaurant menu for a specific nutritional detail or identifying a disclaimer in a lengthy newspaper article tasks that demand careful attention to small but significant details within a broader narrative, akin to Finding Needles in Images (NiM). To address this gap, we introduce NiM, a carefully curated benchmark spanning diverse real-world documents including newspapers, menus, and lecture images, specifically designed to evaluate MLLMs' capability in these intricate tasks. Building on this, we further propose Spot-IT, a simple yet effective approach that enhances MLLMs capability through intelligent patch selection and Gaussian attention, motivated from how humans zoom and focus when searching documents. Our extensive experiments reveal both the capabilities and limitations of current MLLMs in handling fine-grained document understanding tasks, while demonstrating the effectiveness of our approach. Spot-IT achieves significant improvements over baseline methods, particularly in scenarios requiring precise detail extraction from complex layouts.

文档理解多模态细节定位视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。