提升古籍图形检索速度与小图识别精度,10倍提速且准确率创新高。
iDocV2: Leveraging Self-Supervision and Open-Set Detection for Improving Pattern Spotting in Historical Documents

- 用自监督学习优化编码器,结合开放集检测加速搜索
- 小非正方形查询精度达0.612,超越现有最佳水平
- 引入非极大值抑制减少误检,适合古籍数字化项目
随着数字图书的海量增长,通过图形模式快速检索文档变得至关重要。当前针对历史文档的检索与模式识别方法仍存在不足:最先进模型在模式识别上的总体精确率为0.494,小非正方形查询精确率仅为0.427,且处理时间过长,在DocExplore数据集上单次搜索需达7秒,源于其密集策略。为此,本文提出基于改进编码器(iDoc)并采用自监督训练的新模型,结合开放集检测以加速搜索。所提方法在保持与最先进水平相当的检索性能基础上,实现10倍提速;尤其在小非正方形查询上取得新最优结果,精确率达到0.612。相较于前版,本工作引入非极大值抑制机制,有效降低误报率。
原文摘要 · Abstract (English)
Considering the imminent massification of digital books, it has become critical to facilitate searching collections through graphical patterns. Current strategies for document retrieval and pattern spotting in historical documents still need to be improved. State-of-the-art strategies achieve an overall precision of $0.494$ for pattern spotting, where the precision for small non-square queries reaches 0.427. In addition, the processing time is excessive, requiring up to 7 seconds for searching in the DocExplore dataset due to a dense-based strategy used by SOTA models. Therefore, we propose a new model based on a better encoder (iDoc), trained under a self-supervised strategy, and an open-set detector to accelerate searching. Our model achieves competitive results with state-of-the-art pattern spotting and document retrieval, improving speed by 10x. Furthermore, our model reaches a new SOTA performance on the small non-square queries, achieving a new precision of 0.612.Different from the previous version, this leverages non-maximum suppression to reduce false positives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。