ParsiPy让古波斯语文本分析更高效,支持分词、词性标注等核心NLP功能。
ParsiPy: NLP Toolkit for Historical Persian Texts in Python
- 提供分词、词性标注、词形还原等模块,适配古波斯语复杂拼写
- 支持音素转写与词嵌入,提升历史文本数字化处理能力
- 适合研究古波斯语及其它缺乏数字资源的历史语言学者
历史语言研究面临拼写复杂、文本残缺及缺乏标准化数字表示等挑战。本文提出ParsiPy,一个专为历史波斯语设计的Python自然语言处理工具包,包含分词、词形还原、词性标注、音素到转写转换及词嵌入等模块。通过处理帕尔西克语(中古波斯语)文本,验证了该工具在计算文献学中的实用性,推动古代文本的数字化分析与保存。本工作为历史语言的计算研究提供了可复用工具。
原文摘要 · Abstract (English)
The study of historical languages presents unique challenges due to their complex orthographic systems, fragmentary textual evidence, and the absence of standardized digital representations of text in those languages. Tackling these challenges needs special NLP digital tools to handle phonetic transcriptions and analyze ancient texts. This work introduces ParsiPy, an NLP toolkit designed to facilitate the analysis of historical Persian languages by offering modules for tokenization, lemmatization, part-of-speech tagging, phoneme-to-transliteration conversion, and word embedding. We demonstrate the utility of our toolkit through the processing of Parsig (Middle Persian) texts, highlighting its potential for expanding computational methods in the study of historical languages. Through this work, we contribute to computational philology, offering tools that can be adapted for the broader study of ancient texts and their digital preservation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。