首个真实古希腊文错字数据集,助力学者精准识别文本传抄错误。
An Annotated Dataset of Errors in Premodern Greek and Baselines for Detecting Them
- 用BERT条件概率筛选易错词,专家标注1000个真实错字
- 新检测器准确率提升5%,对抄写错误识别更难于印刷或数字化错误
- 适合数字人文、古典文本修复及自然语言处理研究者
随着古代文本在数百年传抄过程中不断流传,错误不可避免地累积。这些错误往往难以发现,正因它们隐蔽已久。以往研究多基于人工生成的错误进行评估,本文首次构建了真实存在的古希腊文错字数据集,使错误检测方法可在真正历史积累的错误上进行评估。通过基于BERT条件概率的指标,我们筛选出1000个高风险词,由领域专家标注为错误或非错误。在此基础上提出并测试新的检测方法,发现基于判别器的检测器表现最佳,对真实错误的真正阳性率提升5%。同时观察到,抄写错误比印刷或数字化错误更难检测。该数据集首次为古代文本真实错误的检测提供了基准,有助于开发更有效的算法,辅助学者修复古籍。
原文摘要 · Abstract (English)
As premodern texts are passed down over centuries, errors inevitably accrue. These errors can be challenging to identify, as some have survived undetected for so long precisely because they are so elusive. While prior work has evaluated error detection methods on artificially-generated errors, we introduce the first dataset of real errors in premodern Greek, enabling the evaluation of error detection methods on errors that genuinely accumulated at some stage in the centuries-long copying process. To create this dataset, we use metrics derived from BERT conditionals to sample 1,000 words more likely to contain errors, which are then annotated and labeled by a domain expert as errors or not. We then propose and evaluate new error detection methods and find that our discriminator-based detector outperforms all other methods, improving the true positive rate for classifying real errors by 5%. We additionally observe that scribal errors are more difficult to detect than print or digitization errors. Our dataset enables the evaluation of error detection methods on real errors in premodern texts for the first time, providing a benchmark for developing more effective error detection algorithms to assist scholars in restoring premodern works.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。