arXiv:2507.20934cs.CV2025-07中稿 · and presented as a…

用文本生成图像解决古籍检索中无样本查询难题

Exploring text-to-image generation for historical document image retrieval

  • 用文本生成图像模拟查询样例,替代传统需真实样本的检索方式
  • 在HisIR19数据集上达到可接受的检索效果,验证方法可行性
  • 适合缺乏查询样本的古籍、档案等独特文献的智能检索场景

基于属性的文档图像检索(ABDIR)是替代传统以图搜图(QBE)的一种新范式,克服了后者需已有查询样本的局限。ABDIR通过卷积神经网络(CNN)分类器提取文档的视觉属性组合来描述内容。本文探索利用文本到图像(T2I)生成技术弥合QBE与ABDIR之间的鸿沟,聚焦于视觉特征多样且独特的古籍文献。提出T2I-QBE方法:使用Leonardo.Ai生成器,结合目标文档类型粗略描述与一组类似ABDIR的视觉属性提示,生成查询图像;再以传统QBE流程比对特征,实现检索。在历史文档数据集HisIR19上的实验验证了该方法的有效性,证明T2I-QBE可用于古籍图像检索。据作者所知,这是首次将T2I生成应用于文档图像检索的研究。

原文摘要 · Abstract (English)

Attribute-based document image retrieval (ABDIR) was recently proposed as an alternative to query-by-example (QBE) searches, the dominant document image retrieval (DIR) paradigm. One drawback of QBE searches is that they require sample query documents on hand that may not be available. ABDIR aims to offer users a flexible way to retrieve document images based on memorable visual features of document contents, describing document images with combinations of visual attributes determined via convolutional neural network (CNN)-based binary classifiers. We present an exploratory study of the use of generative AI to bridge the gap between QBE and ABDIR, focusing on historical documents as a use case for their diversity and uniqueness in visual features. We hypothesize that text-to-image (T2I) generation can be leveraged to create query document images using text prompts based on ABDIR-like attributes. We propose T2I-QBE, which uses Leonardo.Ai as the T2I generator with prompts that include a rough description of the desired document type and a list of the desired ABDIR-style attributes. This creates query images that are then used within the traditional QBE paradigm, which compares CNN-extracted query features to those of the document images in the dataset to retrieve the most relevant documents. Experiments on the HisIR19 dataset of historical documents confirm our hypothesis and suggest that T2I-QBE is a viable option for historical document image retrieval. To the authors' knowledge, this is the first attempt at utilizing T2I generation for DIR.

文本生成图像古籍检索视觉属性AI生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。