arXiv:2608.18437cs.CL2026-08

用少量标注数据和古籍词典实现濒危语言的精准分词

Tangut Word Segmentation under Extreme Resource Scarcity: Integrating Traditional Lexicons and Unlabeled Text

论文配图:Tangut Word Segmentation under Extreme Resource Scarcity: Integrating Traditional Lexicons and Unlabeled Text
图 1 · 摘自论文原文
  • 融合古籍词典与无标注文本,构建可靠性校准的词典图谱
  • 在31,893个词元上达到0.911的平均F1,提升召回率超出标注词汇范围
  • 适合对古文字处理、低资源自然语言处理感兴趣的研究者

Tangut是一种已灭绝的语言,其书写系统不明确标记词边界。本文首次系统研究Tangut分词,基于2,750个专家标注的分段(共31,893个词元)、传统词典和无标注文本。框架结合可靠性校准的词典图谱表示、显式分布统计信息及轻量级字符编码器(通过MLM预训练)。分段级别五折交叉验证显示,词典与统计特征使CRF F1提升至约0.91。完整TangutEncoder达到最高均值F1(0.911),并显著提升召回率,超越标注训练词汇范围。结果表明模型在主题多样的未见篇章中具备泛化能力,文档级迁移效果尚待评估。

原文摘要 · Abstract (English)

Tangut is an extinct language whose script does not explicitly mark word boundaries. We present the first systematic study of Tangut word segmentation using 2,750 expert-annotated segments(31,893 tokens), traditional lexicons, and unlabeled text. Our framework combines a reliability-calibrated lexicon-lattice representation, explicit distributional statistics, and a lightweight character encoder pretrained with MLM. Segment-level five-fold cross-validation shows that lexical and statistical features raise CRF F1 to approximately 0.91. The full TangutEncoder reaches the highest mean F1 (0.911) and improves recall beyond the labeled training vocabulary. These results demonstrate generalization beyond the limited supervised vocabulary across thematically diverse held-out passages, while document-level transfer remains to be evaluated.

古文字处理低资源分词词典融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。