解决历史土耳其语同义词识别难题,提升古籍文本处理效果
Detecting Turkish Synonyms Used in Different Time Periods
- 用正交普鲁斯特法对齐不同时期的词向量空间
- 结合词频相关性,使同义词检测准确率优于基线方法
- 在1960-1980年代表现稳定,后续略有下降
语言动态变化给历史文本的自然语言处理带来挑战,导致下游任务性能下降。土耳其语因20世纪的语言改革呈现快速演变。本文提出两种检测不同时期同义词的方法:第一种使用正交普鲁斯特法对齐不同年代文档生成的嵌入空间;第二种在第一种基础上引入词语年份间频率的斯皮尔曼相关性。实验表明,所提方法优于基线,且在目标时期从1960年代到1980年代时表现一致;但在后续时期性能略有下降。
原文摘要 · Abstract (English)
Dynamic structure of languages poses significant challenges in applying natural language processing models on historical texts, causing decreased performance in various downstream tasks. Turkish is a prominent example of rapid linguistic transformation due to the language reform in the 20th century. In this paper, we propose two methods for detecting synonyms used in different time periods, focusing on Turkish. In our first method, we use Orthogonal Procrustes method to align the embedding spaces created using documents written in the corresponding time periods. In our second method, we extend the first one by incorporating Spearman's correlation between frequencies of words throughout the years. In our experiments, we show that our proposed methods outperform the baseline method. Furthermore, we observe that the efficacy of our methods remains consistent when the target time period shifts from the 1960s to the 1980s. However, their performance slightly decreases for subsequent time periods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。