用细粒度分析法发现甲骨文新重复样本,提升考古辨识效率。
Explainable Coarse-to-Fine Ancient Manuscript Duplicates Discovery
- 分步融合低层特征与高层语义匹配,提升识别准确率。
- 在真实场景中发现60余对专家遗漏的甲骨文重复内容。
- 模型兼具可解释性与高效计算,适合考古与历史研究者使用。
古代手稿是古代语言语料的主要来源,但因无意重复出版或故意伪造,常出现重复现象。例如死海古卷包含伪造残片,而甲骨文(OB)则既有重印材料也存在伪造品。识别甲骨文重复内容对考古整理与古代史研究具有重要意义。本文设计了一种渐进式甲骨文重复发现框架,结合无监督底层关键点匹配与高层文本中心内容匹配,以语义感知和可解释性精炼并排序候选重复项。相较于当前最先进的基于内容的图像检索与图像匹配方法,本模型在召回率上表现相当,且在Top-5与Top-15检索结果中取得最高简化平均倒数排名分数,并实现显著加速的计算效率。在实际部署中,已发现超过60对此前被领域专家遗漏的甲骨文重复样本。代码、模型与真实结果详见:https://github.com/cszhangLMU/OBD-Finder/。
原文摘要 · Abstract (English)
Ancient manuscripts are the primary source of ancient linguistic corpora. However, many ancient manuscripts exhibit duplications due to unintentional repeated publication or deliberate forgery. The Dead Sea Scrolls, for example, include counterfeit fragments, whereas Oracle Bones (OB) contain both republished materials and fabricated specimens. Identifying ancient manuscript duplicates is of great significance for both archaeological curation and ancient history study. In this work, we design a progressive OB duplicate discovery framework that combines unsupervised low-level keypoints matching with high-level text-centric content-based matching to refine and rank the candidate OB duplicates with semantic awareness and interpretability. We compare our model with state-of-the-art content-based image retrieval and image matching methods, showing that our model yields comparable recall performance and the highest simplified mean reciprocal rank scores for both Top-5 and Top-15 retrieval results, and with significantly accelerated computation efficiency. We have discovered over 60 pairs of new OB duplicates in real-world deployment, which were missed by domain experts for decades. Code, model and real-world results are available at: https://github.com/cszhangLMU/OBD-Finder/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。