用CLIP模型实现1360万份17世纪西属美洲公证档案的图文检索
SARCLIP: A Scalable CLIP-Based Retrieval System for Seventeenth-Century Spanish American Notary Records

- 基于专家标注数据微调CLIP模型,实现手写体图文匹配
- 在近100卷微缩胶片上完成超大规模近似最近邻检索
- 支持可视化浏览、标注与模型持续迭代,形成人机闭环
历史手稿档案因书写不一、古体拼写和缺乏可靠转录而难以进行标准文本搜索。我们提出SARCLIP(西班牙美洲公证记录遇见CLIP),一个部署于阿根廷国家档案馆的17世纪西班牙美洲公证记录检索系统,该语料库包含超过1360万份词图像块,覆盖100多卷微缩胶片("rollos")。SARCLIP基于对比微调的CLIP ViT-B/16模型,通过FAISS索引实现近完整语料库的近似最近邻检索,并结合伪相关反馈(Rocchio)优化前k个结果,同时建立人机闭环:支持可视化文档浏览、画布式图像块标注及定期模型再训练。相较于初期仅在5卷子集评估的研究原型,SARCLIP已作为完整交互式工具运行于近完整语料库,演示中观众可实时体验完整搜索、浏览、标注与重训练流程。
原文摘要 · Abstract (English)
Historical manuscript archives resist standard text search due to inconsistent handwriting, archaic orthography, and the absence of reliable transcriptions at scale. We present SARCLIP (Spanish American Notary Records Meets CLIP), a deployed retrieval system for the National Archives of Argentina's seventeenth-century Spanish American notary records, a corpus of more than 13.6 million word-image patches spanning over 100 microfilm rolls ("rollos"). SARCLIP is built on a CLIP ViT-B/16 model contrastively fine-tuned on paleography-expert-annotated data, and extends prior work by (1) scaling approximate nearest-neighbor retrieval to the near-complete corpus via a FAISS index, (2) refining top-k results through pseudo-relevance feedback (Rocchio), and (3) closing a human-in-the-loop cycle through visual document browsing, canvas-based patch annotation, and periodic model retraining. Unlike the system's initial research prototype, which evaluated retrieval on a small five-rollo subset, SARCLIP is demonstrated as a complete, interactive tool operating over the near-complete corpus. Attendees experience the full search, browse, annotate, and retrain workflow live during this demonstration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。