arXiv:2501.04828cs.CL2025-01被引 2

首个历史土耳其语NLP资源与模型,助力古籍文本分析

Building Foundations for Natural Language Processing of Historical Turkish: Resources and Models

  • 构建首个历史土耳其语命名实体数据集与句法树库
  • 在命名实体识别等任务上达90.29% F1,表现优异
  • 适合历史语言学、数字人文研究者使用

本文为历史土耳其语自然语言处理建立了基础资源与模型,该领域在计算语言学中长期被忽视。我们首次推出命名实体识别数据集HisTR和通用依存树库OTA-BOUN,以及基于这些数据训练的Transformer模型,涵盖命名实体识别、依存句法分析和词性标注任务。同时发布奥斯曼文转写语料库OTC,覆盖广泛历史时期。实验结果表明,模型在理解历史语言结构的任务中表现突出:命名实体识别达90.29% F1,依存句法分析LAS为73.79%,词性标注达94.98% F1。研究也揭示了领域适应与不同时期语言差异等挑战。所有资源与模型均已公开于https://hf.co/bucolin,可作为未来研究基准。

原文摘要 · Abstract (English)

This paper introduces foundational resources and models for natural language processing (NLP) of historical Turkish, a domain that has remained underexplored in computational linguistics. We present the first named entity recognition (NER) dataset, HisTR, and the first Universal Dependencies treebank, OTA-BOUN, for a historical form of the Turkish language along with transformer-based models trained using these datasets for named entity recognition, dependency parsing, and part-of-speech tagging tasks. Furthermore, we introduce the Ottoman Text Corpus (OTC), a clean corpus of transliterated historical Turkish texts that spans a wide range of historical periods. Our experimental results demonstrate prominent improvements in the computational analysis of historical Turkish, achieving strong performance on tasks that require understanding of historical linguistic structures -- specifically, 90.29% F1 in named entity recognition, 73.79% LAS for dependency parsing, and 94.98% F1 for part-of-speech tagging. They also highlight existing challenges, such as domain adaptation and language variations between time periods. All the resources and models presented are available at https://hf.co/bucolin to serve as a benchmark for future progress in historical Turkish NLP.

历史语言土耳其语NER依存句法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。