用内部语言特征自动识别借词,不依赖外部信息。
Feature-Refined Unsupervised Model for Loanword Detection
- 仅用语言内部特征提取并迭代优化借词判断
- 在六种印欧语系语言上表现优于基线方法
- 适合历史语言学和跨语言研究者使用
我们提出一种无监督方法用于检测借词,即从一种语言借用到另一种语言的词汇。以往研究多依赖语言外部信息,可能引入循环性与工作流程限制。本文模型仅基于语言内部信息,处理单语与多语词表中的原生词与借词。通过提取关键语言特征、打分并概率映射,迭代优化初始结果,识别并归纳新出现模式直至收敛。该混合方法结合语言学与统计线索引导发现过程。我们在六种标准印欧语系语言(英语、德语、法语、意大利语、西班牙语、葡萄牙语)的数据集上评估,实验表明模型性能优于基线方法,尤其在跨语言数据扩展时提升显著。
原文摘要 · Abstract (English)
We propose an unsupervised method for detecting loanwords i.e., words borrowed from one language into another. While prior work has primarily relied on language-external information to identify loanwords, such approaches can introduce circularity and constraints into the historical linguistics workflow. In contrast, our model relies solely on language-internal information to process both native and borrowed words in monolingual and multilingual wordlists. By extracting pertinent linguistic features, scoring them, and mapping them probabilistically, we iteratively refine initial results by identifying and generalizing from emerging patterns until convergence. This hybrid approach leverages both linguistic and statistical cues to guide the discovery process. We evaluate our method on the task of isolating loanwords in datasets from six standard Indo-European languages: English, German, French, Italian, Spanish, and Portuguese. Experimental results demonstrate that our model outperforms baseline methods, with strong performance gains observed when scaling to cross-linguistic data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。