为俄国航天先驱的手稿建立可机器读取的目录与转录数据集
A machine-readable catalogue of the Tsiolkovsky papers (fond 555, Archive of the Russian Academy of Sciences), and a way to measure how well its handwriting can be read
- 构建包含2019份文件、51008页扫描件的可检索目录与全文转录
- 通过打字稿与手稿比对,测出手写识别平均准确率中位数为37%
- 发现手稿修订版本间仅19%词汇重合,无法逐字比对
康斯坦丁·齐奥尔科夫斯基(1857-1935)个人档案存于俄罗斯科学院档案馆第555号典藏。尽管该档案已扫描并发布图像,但缺乏可查询目录、全文搜索功能和数据集,仅能逐页浏览。本文构建了包含2,019份文件、51,008页扫描件的机器可读目录,为1,969份文件提供日期标注,并对每页进行手写体与打印体分类;同时完成全库机器转录。提出一种在无真实标注情况下评估手写识别准确性的方法:利用打字稿与手稿双重记录,通过对比同一内容的两种版本分离出阅读错误。在224份文件共1,759对样本中,两轮手写识别中位数仅37%词汇一致,最长连续正确词串为10个;这一结果在样本量扩大六倍后保持稳定。对两份有正式出版版本的文件验证,估计值与真实情况偏差小于1个百分点,等级相关系数达0.92,表明该方法可靠。然而,同一作品不同版本间仅19%词汇重合,低于单页识别一致性水平,因此无法实现逐字校勘。该限制被明确报告并纳入工具设计。
原文摘要 · Abstract (English)
The personal archive of Konstantin Tsiolkovsky (1857-1935) is held as fond 555 of the Archive of the Russian Academy of Sciences. The archive scanned the fond and published the images, but with no queryable catalogue, no full-text search and no dataset: the holdings can only be browsed one page at a time. This paper describes a machine-readable catalogue of all 2,019 files and 51,008 scans, a dating for 1,969 files taken from the archive's own descriptions, a page-level classification of every scan into handwriting and typescript, and a machine transcription of the fond in full: all 2,019 files and all 51,008 scans. It also reports a way to measure handwritten-text-recognition accuracy in an archive with no ground truth. Archives of the typewriter era often preserve one text twice, as manuscript and as a typed copy; transcribing both and comparing isolates the reading error, since source and pipeline are identical and only page difficulty differs. Across 1,759 such pairs from 224 files, two readings of a handwritten page agree on a median 37% of words and share a longest verbatim run of a median 10 words; the median is unchanged from the 294 pairs of the first version, on a sample six times larger. On two files that also have a published edition the estimate can be checked against ground truth: it is unbiased to within a percentage point and ranks pages as the truth does (rank correlation 0.92 where the edition is a faithful witness). This bounds use: two variants of one work here share 19% of words, below the rate at which two readings of a single page agree, so the redactions cannot be collated word by word at this quality. That negative result is reported as such, and the constraint is built into the tool.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。