arXiv:2512.20174cs.CVcs.CL2025-12CVPR被引 6

用自然语言查询实现精准文档图像检索,解决细粒度语义匹配难题。

Towards Natural Language-Based Document Image Retrieval: New Dataset and Benchmark

  • 以大模型生成+人工验证的5万条细粒度文本查询构建新数据集
  • 零样本与微调实验验证主流视觉语言模型在跨模态检索中的性能
  • 提出两阶段高效检索方法,兼顾准确率与计算效率,适合实际应用

文档图像检索(DIR)旨在根据查询从图库中检索文档图像。现有方法主要依赖图像查询,仅能处理粗粒度类别如报纸或收据,难以应对真实场景中常见的细粒度文本查询。为此,我们提出自然语言文档图像检索(NL-DIR)基准及评估指标,采用自然语言描述作为语义丰富的查询。NL-DIR数据集包含41,000张真实文档图像,每张图像配以五条由大语言模型生成并经人工验证的高质量细粒度语义查询。我们对主流对比视觉语言模型和无OCR视觉文档理解(VDU)模型进行零样本与微调评估,并探索一种两阶段检索方法,在提升性能的同时实现时间和空间效率优化。我们希望该基准能为VDU领域带来新机遇。数据集与代码将公开于 huggingface.co/datasets/nianbing/NL-DIR。

原文摘要 · Abstract (English)

Document image retrieval (DIR) aims to retrieve document images from a gallery according to a given query. Existing DIR methods are primarily based on image queries that retrieve documents within the same coarse semantic category, e.g., newspapers or receipts. However, these methods struggle to effectively retrieve document images in real-world scenarios where textual queries with fine-grained semantics are usually provided. To bridge this gap, we introduce a new Natural Language-based Document Image Retrieval (NL-DIR) benchmark with corresponding evaluation metrics. In this work, natural language descriptions serve as semantically rich queries for the DIR task. The NL-DIR dataset contains 41K authentic document images, each paired with five high-quality, fine-grained semantic queries generated and evaluated through large language models in conjunction with manual verification. We perform zero-shot and fine-tuning evaluations of existing mainstream contrastive vision-language models and OCR-free visual document understanding (VDU) models. A two-stage retrieval method is further investigated for performance improvement while achieving both time and space efficiency. We hope the proposed NL-DIR benchmark can bring new opportunities and facilitate research for the VDU community. Datasets and codes will be publicly available at huggingface.co/datasets/nianbing/NL-DIR.

文档检索自然语言查询视觉语言模型细粒度匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。