arXiv:2608.22214cs.AIcs.MM2026-08

提出图文联合提取新任务,精准定位查询对应的文本与图像。

Query-Driven Multimodal Information Extraction from Long Documents

论文配图:Query-Driven Multimodal Information Extraction from Long Documents
图 1 · 摘自论文原文
  • 构建双层分类体系,理解查询意图与文档内容
  • 在910个实例上,多智能体框架显著优于单模型
  • 适合需要跨模态精确检索的领域文档应用

在特定领域的长篇多模态文档中,图像与文本共同传递复杂知识,仅靠纯文本难以完整捕捉。现有范式如DocVQA主要生成文本答案或定位证据区域,无法输出查询相关的具体文本属性值及其对应图像。为此,我们提出查询驱动的图文联合提取任务,要求模型输出查询所需的文本属性值及对应图像边界框。针对用户意图与文档内容双重挑战,设计了查询与实例两级分类体系。进一步构建了首个高质量人工标注基准ITJoint,包含2,455页领域文档、316个查询和910个答案实例,涵盖大量非装饰性图像。评估多个独立视觉语言模型后,提出Q2IT多智能体协作框架,包含证据收集、页面选择与目标图像定位三阶段协同代理。采用联合评估策略同时衡量文本提取与图像定位性能,实验表明独立VLM表现有限,而Q2IT显著提升,但仍存在明显差距。

原文摘要 · Abstract (English)

In domain-specific multimodal long documents, images and text jointly convey complex knowledge that cannot be fully captured by plain text alone. However, existing paradigms like DocVQA primarily focus on generating textual answers or localizing evidence regions, rather than outputting query-specific textual attribute values and corresponding images. To address this gap, we propose query-driven image-text joint extraction from long documents, requiring models to output query-requested textual attribute values and corresponding image bounding boxes. Based on challenges related to both user intent and document content, we designed a two-level taxonomy that operates at the query and instance levels. Further, we construct ITJoint, the first high-quality, manually annotated benchmark for this new task, comprising 2,455 pages of domain-specific documents with numerous non-decorative images, 316 queries, and 910 answer instances. Finally, we evaluate representative standalone Vision-Language Models from different providers and further design Q2IT, a multi-agent collaborative framework consisting of three progressively collaborating agents for evidence collection, page selection, and target-image localization. Using a joint evaluation approach that assesses both text extraction and image localization, our experiments show that standalone VLMs struggle with this task, while Q2IT significantly improves performance on ITJoint, although a substantial gap remains toward perfect results.

多模态信息提取图文对齐智能代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。