15世纪纽伦堡信件手稿多转录数据集,助力人文研究的文本分析
Nuremberg Letterbooks: A Multi-Transcriptional Dataset of Early 15th Century Manuscripts for Document Analysis
- 提供三种转录类型,支持手写体识别与作者辨识
- 含1711页、10位抄写员,每页可标注不同作者
- 适合作为人文领域文档分析的基准数据集
当前文档分析领域的多数数据集采用高度标准化标签,虽简化特定任务,但结果难以直接用于人文学科研究。为此,本文推出纽伦堡信件手稿数据集(Nuremberg Letterbooks),包含15世纪早期的4本历史手稿,共1711页,由10位抄写员书写。数据集提供三种手写文本识别转录:基础型、外交型和规范化型,其中后两种均提供含缩写扩展与不含扩展的版本。通过字母编号与抄写员编号结合,支持跨页作者识别。技术验证中建立了各项任务基线,证明数据一致性,并为后续研究提供可复现基准。
原文摘要 · Abstract (English)
Most datasets in the field of document analysis utilize highly standardized labels, which, while simplifying specific tasks, often produce outputs that are not directly applicable to humanities research. In contrast, the Nuremberg Letterbooks dataset, which comprises historical documents from the early 15th century, addresses this gap by providing multiple types of transcriptions and accompanying metadata. This approach allows for developing methods that are more closely aligned with the needs of the humanities. The dataset includes 4 books containing 1711 labeled pages written by 10 scribes. Three types of transcriptions are provided for handwritten text recognition: Basic, diplomatic, and regularized. For the latter two, versions with and without expanded abbreviations are also available. A combination of letter ID and writer ID supports writer identification due to changing writers within pages. In the technical validation, we established baselines for various tasks, demonstrating data consistency and providing benchmarks for future research to build upon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。