构建可搜索的政府文件系统,支持图文混合查询。
GovScape: A Public Multimodal Search System for 70 Million Pages of Government PDFs
- 基于1000万份政府PDF建立多模态搜索系统
- 实现语义搜索与图像内容检索,如查找红印文档或饼图
- 仅耗资1500美元完成预处理,每美元处理4.7万页
过去三十年的努力已形成包含数十亿网页快照和数百PB数据的网络档案。仅2020年任期结束网络归档就包含数百万份联邦政府生成的PDF文件。尽管数据保存成功,但访问与发现仍面临挑战。目前对这些文件的浏览方式仅限于下载、逐份查看及基础关键词搜索。本文介绍GovScape,一个面向10,015,993份联邦政府PDF(共70,958,487页)的公开多模态搜索系统——据我们所知,这是2020年爬取中所有不超过50页的可渲染PDF。GovScape支持四种主要搜索方式:(1)基于领域、爬取日期等元数据的筛选;(2)对PDF文本的精确匹配搜索;(3)语义文本搜索;(4)针对单页内容的视觉搜索,使用户能提出如“红印文件”或“饼图”等结构化查询。文章详述了系统的构成组件,包括搜索功能、嵌入流水线、系统架构与开源代码库。显著的是,该系统对1000万份PDF的预处理计算成本约1500美元,相当于每美元处理47,000页,展现了快速扩展的潜力。我们已着手推进在1亿+份PDF规模下的多模态搜索。系统网址:https://www.govscape.net。
原文摘要 · Abstract (English)
Efforts over the past three decades have produced web archives containing billions of webpage snapshots and petabytes of data. The End of Term Web Archive alone contains, among other file types, millions of PDFs produced by the federal government. While preservation with web archives has been successful, significant challenges for access and discoverability remain. For example, current affordances for browsing the End of Term PDFs are limited to downloading and browsing individual PDFs, as well as performing basic keyword search across them. In this paper, we introduce GovScape, a public search system that supports multimodal searches across 10,015,993 federal government PDFs from the 2020 End of Term crawl (70,958,487 total PDF pages) - to our knowledge, all renderable PDFs in the 2020 crawl that are 50 pages or under. GovScape supports four primary forms of search over these 10 million PDFs: in addition to providing (1) filter conditions over metadata facets including domain and crawl date and (2) exact text search against the PDF text, we provide (3) semantic text search and (4) visual search against the PDFs across individual pages, enabling users to structure queries such as "redacted documents" or "pie charts." We detail the constituent components of GovScape, including the search affordances, embedding pipeline, system architecture, and open source codebase. Significantly, the total estimated compute cost for GovScape's pre-processing pipeline for 10 million PDFs was approximately $1,500, equivalent to 47,000 PDF pages per dollar spent on compute, demonstrating the potential for immediate scalability. Accordingly, we outline steps that we have already begun pursuing toward multimodal search at the 100+ million PDF scale. GovScape can be found at https://www.govscape.net.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。