arXiv:2511.11552cs.CVcs.CL2025-11被引 14

DocLens用多智能体工具框架提升长图文文档的理解精度。

DocLens : A Tool-Augmented Multi-Agent Framework for Long Visual Document Understanding

  • 通过多智能体协作实现从全文到关键视觉元素的精准定位。
  • 在MMLongBench-Doc和FinRAGBench-V上超越人类专家表现。
  • 适合处理视觉主导或无法回答的问题,减少模型幻觉。

理解信息分散在大量文字与视觉元素中的长篇视觉文档,是现代视觉语言模型(VLMs)面临的关键挑战。现有方法在证据定位上表现不佳,难以检索相关页面并忽略视觉元素中的细粒度信息,导致性能受限且易产生幻觉。为此,我们提出DocLens——一种工具增强的多智能体框架,可像镜头一样聚焦证据。它先从全文导航至相关页面上的特定视觉元素,再通过采样-裁决机制生成单一可靠答案。结合Gemini-2.5-Pro,DocLens在MMLongBench-Doc和FinRAGBench-V上达到当前最佳性能,甚至超越人类专家。其优势在以视觉为中心及无解查询中尤为显著,体现了更强的定位能力。

原文摘要 · Abstract (English)

Comprehending long visual documents, where information is distributed across extensive pages of text and visual elements, is a critical but challenging task for modern Vision-Language Models (VLMs). Existing approaches falter on a fundamental challenge: evidence localization. They struggle to retrieve relevant pages and overlook fine-grained details within visual elements, leading to limited performance and model hallucination. To address this, we propose DocLens, a tool-augmented multi-agent framework that effectively ``zooms in'' on evidence like a lens. It first navigates from the full document to specific visual elements on relevant pages, then employs a sampling-adjudication mechanism to generate a single, reliable answer. Paired with Gemini-2.5-Pro, DocLens achieves state-of-the-art performance on MMLongBench-Doc and FinRAGBench-V, surpassing even human experts. The framework's superiority is particularly evident on vision-centric and unanswerable queries, demonstrating the power of its enhanced localization capabilities.

文档理解多智能体视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。