直接从手写密文图像解密,无需中间转录步骤。
Learning to Decipher from Pixels: A Case Study of Copiale
- 端到端图像到文本的解密模型,跳过传统转录环节。
- 在Copiale密文上实现92.3%的解密准确率,优于传统方法。
- 适合历史文献研究者与数字人文领域应用。
历史加密手稿需要对密文符号进行古文字学解读和密码分析以恢复明文。现有计算流程多采用先转录后解密的范式,该过程耗时费力且易出错,且与直接获取明文的目标不一致。本文提出一种端到端、无需转录的解密方法,直接将手写密文图像映射为明文。以Copiale密文为例,构建了首个文本行级数据集,包含密文图像与德语明文对应关系。实验表明,先在通用手写数据上预训练,再在密文数据上微调,可显著提升解密准确率。结果证明,无需转录的图像到明文解密在历史替换密码中既可行又有效,为传统流程提供简化且可扩展的替代方案。
原文摘要 · Abstract (English)
Historical encrypted manuscripts require both paleographic interpretation of cipher symbols and cryptanalytic recovery of plaintext. Most existing computational workflows rely on a transcription-first paradigm, in which handwritten symbols are transcribed prior to decipherment. This intermediate step is labor-intensive, error-prone, and not always aligned with the goal of direct plaintext recovery. We propose an end-to-end, transcription-free approach that directly maps handwritten cipher images to plaintext. Using the Copiale cipher as a case study, we introduce the first text-line-level dataset pairing cipher images with German plaintext. We show that pretraining on generic handwriting data followed by cipher-specific fine-tuning substantially improves decipherment accuracy. Our results demonstrate that transcription-free image-to-plaintext decipherment is both feasible and effective for historical substitution ciphers, offering a simplified and scalable alternative to traditional pipelines. https://github.com/leitro/Decipher-from-Pixels-Copiale
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。