构建首个中世纪到近代手稿的开源语料库,支持文字识别与实体识别研究
TRIDIS: A Comprehensive Medieval and Early Modern Corpus for HTR and NER
- 整合多个开放资源,统一标注规则与元数据
- 提出基于嵌入空间异常检测的跨域测试集划分方法
- 提供TrOCR和MiniCPM2.5基线,助力手写文本分析
本文介绍TRIDIS(Tria Digita Scribunt),一个面向中世纪至近代手稿的开源语料库。该语料库汇集多个已公开的遗留数据集(均采用开放许可),并包含丰富的元数据描述。尽管此前已有部分研究引用其内容,但本工作首次提供全面统一的概述,并强化其构成分析。我们详细说明:(i) 各主要子语料库的叙事、时间与编辑背景;(ii) 半抄本式转录规则(包括扩展、规范化与标点);(iii) 基于联合嵌入空间异常检测的挑战性跨域测试集划分策略;(iv) 使用TrOCR和MiniCPM2.5进行的初步基线实验,对比随机与异常值驱动的测试划分效果。TRIDIS旨在推动中世纪至近代文本遗产在手写文本识别(HTR)与命名实体识别(NER)方面的联合鲁棒研究。
原文摘要 · Abstract (English)
This paper introduces TRIDIS (Tria Digita Scribunt), an open-source corpus of medieval and early modern manuscripts. TRIDIS aggregates multiple legacy collections (all published under open licenses) and incorporates large metadata descriptions. While prior publications referenced some portions of this corpus, here we provide a unified overview with a stronger focus on its constitution. We describe (i) the narrative, chronological, and editorial background of each major sub-corpus, (ii) its semi-diplomatic transcription rules (expansion, normalization, punctuation), (iii) a strategy for challenging out-of-domain test splits driven by outlier detection in a joint embedding space, and (iv) preliminary baseline experiments using TrOCR and MiniCPM2.5 comparing random and outlier-based test partitions. Overall, TRIDIS is designed to stimulate joint robust Handwritten Text Recognition (HTR) and Named Entity Recognition (NER) research across medieval and early modern textual heritage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。