arXiv:2604.01425cs.CL2026-04

用上下文区分印地语同义词的词源,发现语境能捕捉历史来源的细微差异。

The power of context: Random Forest classification of near synonyms. A case study in Modern Hindi

  • 基于词嵌入的随机森林模型分类词源
  • 在语义无关情况下仍能准确区分梵语与波斯语来源
  • 适合语言学、历史语义与自然语言处理研究者

同义现象虽普遍存在却难以解释。理论上绝对同义词不应存在,因其无法扩展语言表达力。然而有观点认为,即便同义词表意相同,也可能反映不同视角或携带不同文化内涵,这一观点极少被量化验证。在印地语中,长期受波斯语影响产生了大量波斯-阿拉伯借词,与对应的梵语词形成众多同义对。本研究探讨这些借词出现数百年后,是否仅凭分布数据即可区分其词源,且不依赖语义内容。训练于印地语同义词词嵌入的随机森林模型成功分类出词源(梵语或波斯-阿拉伯),即使词语语义无关。结果表明使用模式保留了词源痕迹,为语境编码词源信号提供了定量证据。这支持同义词可能反映细微但系统性差异,并暗示词源相关词可能构成不同的概念子空间,形成由历史起源塑造的新类型语义框架。整体说明上下文可捕捉超越传统语义相似性的细微差别。

原文摘要 · Abstract (English)

Synonymy is a widespread yet puzzling linguistic phenomenon. Absolute synonyms theoretically should not exist, as they do not expand language's expressive potential. However, it was suggested that even if synonyms denote the same concept, they may reflect different perspectives or carry distinct cultural associations, claims that have rarely been tested quantitatively. In Hindi, prolonged contact with Persian produced many Perso-Arabic loanwords coexisting with their Sanskrit counterpart, forming numerous synonym pairs. This study investigates whether centuries after these borrowings appeared in the Subcontinent their origin can still be distinguished using distributional data alone and regardless of their semantic content. A Random Forest trained on word embeddings of Hindi synonyms successfully classified words by Sanskrit or Perso-Arabic origin, even when they were semantically unrelated, suggesting that usage patterns preserve traces of etymology. These findings provide quantitative evidence that context encodes etymological signals and that synonymy may reflect subtle but systematic distinctions linked to origin. They support the idea that synonymous words can offer different perspectives and that etymologically related words may form distinct conceptual subspaces, creating a new type of semantic frame shaped by historical origin. Overall, the results highlight the power of context in capturing nuanced distinctions beyond traditional semantic similarity.

语义分析词源识别随机森林印地语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。