arXiv:2609.07159cs.CV2026-09

无需标注数据,自动在古手稿中定位密文符号,准确率远超现有方法。

Unsupervised Domain Adaptation for Symbol Spotting in Historical Encrypted Manuscripts

论文配图:Unsupervised Domain Adaptation for Symbol Spotting in Historical Encrypted Manuscripts
图 1 · 摘自论文原文
  • 用联合编码器+风格迁移,让字体与手稿符号对齐
  • 在14页古籍上比CLIP高0.194的检索准确率
  • 可无监督识别未知手稿的字母体系,适合历史学家使用

破译历史加密手稿是数字人文的核心挑战:在开始转录前,必须先识别并分析底层密码字母表。本文通过符号定位解决此问题:给定一组渲染的字体字形作为候选字母表,任务是在未见的手写文档中判断其字符是否存在及位置,且不依赖目标脚本的标注样本。主要难点在于干净的数字字体查询与退化的手写符号之间存在显著领域差异。我们提出三阶段无监督流程,结合SimCLR+DANN联合编码器生成领域不变的字形表示,并在检索时引入嵌入空间风格适配机制,无需重新训练。在七个加密手稿集合的十四页样本上实验表明,该方法显著优于零样本基础模型(如CLIP和DINOv2),P@1指标比CLIP ViT-L/14高出0.194;同时超越特定任务训练的基线模型0.138。此外,我们在完全无监督设置下计算的Raw-Cover指标,能有效生成脚本族指纹,识别未知文档的底层字母体系,对古文字学家、历史学者等具有直接应用价值。

原文摘要 · Abstract (English)

The decipherment of historical encrypted manuscripts poses a fundamental challenge in Digital Humanities: before any transcription can begin, the symbol inventory of the underlying cipher alphabet must first be identified and characterized. We address this challenge through symbol spotting: given a candidate alphabet specified as a set of rendered font glyphs, the task is to determine whether and where its characters appear in an unseen handwritten document, without any labeled examples from the target script. The main difficulty lies in the domain gap between clean, digitally rendered font queries and degraded handwritten manuscript symbols. We propose a three-stage pipeline that bridges this gap without manual annotation, combining a joint SimCLR+DANN encoder for domain-invariant glyph representations with an embedding-space style-adaptation mechanism applied at retrieval time, requiring no re-training. Experiments on fourteen pages from seven encrypted manuscript collections show that our method outperforms zero-shot foundation models, including CLIP and DINOv2, by a large margin ($+0.194$ P@1 over CLIP ViT-L/14), and surpasses task-specific trained baselines by $+0.138$ P@1. We further demonstrate that the Raw-Cover metric, computed in a fully unsupervised setting, provides a meaningful script-family fingerprint that identifies the underlying alphabet of an unknown document. This capability is of direct practical relevance to palaeographers, historians, and other researchers working with undeciphered manuscripts.

符号定位无监督学习古文书

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。