用文字形态分析实现历史手稿的可解释测量,仅需行级标注
Leveraging Morphology for Historical Script Metrological Analysis

- 基于Transformer和原型重建,从行级文本标注学习字符原型
- 在160页古籍上实现字符、双字词和间距的精准测量
- 只需一列文本数据,适合小样本历史文献研究
手写文本识别进步使历史文档大规模转录成为可能,但仍难以提供可解释的视觉度量以支持古文字学研究。本文核心洞察是:通过形态学分析,特别是从行级转录中学习字符原型,可定义可扩展、有意义且稳定的古文字度量。我们采用基于Transformer的检测架构与基于原型的行重建模块,学习字符原型及其出现、变形与位置分布。贡献有二:一是提出深度架构与学习方法,仅需行级标注即可高效建模字符,显著优于可学习打字机基线,并实现准确的字符边界框预测;二是展示该架构所支持的自动度量在字符、双字词及图形单元间距方面的古文字学相关性。为此,我们扩展了14世纪末由查理五世委托、四人执笔的《巴黎国家图书馆法文2813号抄本》的标注至160页,可视化测量结果,揭示其不仅能区分图形特征,还能发现并分析细微差异。该案例展示了方法的可扩展性与低数据需求——每页仅需一列文本即可完成测量。数据与代码公开于:https://malamatenia.github.io/morphology4metrology-analysis。
原文摘要 · Abstract (English)
Advances in handwritten text recognition have enabled large-scale transcription of historical documents, but still provide limited access to interpretable visual measurements for paleography, the study of historical scripts. In this paper, our main insight is that morphological script analysis, in particular the capacity to learn character prototypes from line-level transcriptions, enables the definition of scalable, meaningful, and stable paleographic measurements. More precisely, we leverage a transformer-based detection architecture together with a prototype-based line reconstruction module to learn prototypical characters and their occurrence, deformation, and positioning. Our contributions are twofold. First, we introduce a deep architecture and learning methodology that enables efficient character modeling with only line-level transcription supervision, significantly improving over the Learnable Typewriter baseline and enabling accurate character bounding box prediction, unlocking its potential for paleographic measurements. Second, we introduce and demonstrate the paleographical relevance of automatic measurements enabled by our architecture for characters, bi-grams, and spaces between graphical units. For this demonstration, we extend the annotations of the codex Paris, BnF, fr. 2813, commissioned in the late fourteenth century by Charles V and copied by four hands, to 160 pages. We visualize our measurements over these pages, showing how they enable us not only to differentiate graphical profiles, but also to discover and analyze subtle variations. This case study outlines the scalability of our approach and its frugality in terms of required training data, since a single column of text is sufficient to compute our measurements on each of the 160 pages. Data and code are publicly available at: https://malamatenia.github.io/morphology4metrology-analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。