arXiv:2503.18226cs.CL2025-03中稿 · NAACL被引 2

用NLP分析《梨俱吠陀》诗篇,发现语言模式与古籍分组高度吻合。

Mapping Hymns and Organizing Concepts in the Rigveda: Quantitatively Connecting the Vedic Suktas

  • 创新改进潜在语义分析(LSA),捕捉诗篇深层语义关联。
  • 新LSA方法识别出7个经典主题组,统计显著且模块度达0.944。
  • 优于SBERT和Doc2Vec,适合古文献文本结构挖掘研究者。

由于《梨俱吠陀》语言古老、结构诗意且文本量庞大,获取其内容并深入理解极具挑战。本研究采用自然语言处理技术,对现代英文译本中1,028篇诗篇(suktas)进行预处理,分别使用三种嵌入方法:本文提出的新型潜在语义分析(LSA)改进版、SBERT和Doc2Vec。通过UMAP降维后,构建诗篇间的k近邻网络,并利用Louvain、Leiden及标签传播算法进行社区检测。结果表明,仅新型LSA结合Leiden方法得出的网络具有统计显著性(z = 2.726, p < .01),模块度为0.944。在分析的七个著名诗篇分组(如创世、丧葬、水等)中,该方法成功识别全部七组;而Doc2Vec未达显著性,无法分离相关诗篇;SBERT虽识别出四组独立主题,但将另三组错误合并成单一混合群组,且整体网络不显著。

原文摘要 · Abstract (English)

Accessing and gaining insight into the Rigveda poses a non-trivial challenge due to its extremely ancient Sanskrit language, poetic structure, and large volume of text. By using NLP techniques, this study identified topics and semantic connections of hymns within the Rigveda that were corroborated by seven well-known groupings of hymns. The 1,028 suktas (hymns) from the modern English translation of the Rigveda by Jamison and Brereton were preprocessed and sukta-level embeddings were obtained using, i) a novel adaptation of LSA, presented herein, ii) SBERT, and iii) Doc2Vec embeddings. Following an UMAP dimension reduction of the vectors, the network of suktas was formed using k-nearest neighbours. Then, community detection of topics in the sukta networks was performed with the Louvain, Leiden, and label propagation methods, whose statistical significance of the formed topics were determined using an appropriate null distribution. Only the novel adaptation of LSA using the Leiden method, had detected sukta topic networks that were significant (z = 2.726, p < .01) with a modularity score of 0.944. Of the seven famous sukta groupings analyzed (e.g., creation, funeral, water, etc.) the LSA derived network was successful in all seven cases, while Doc2Vec was not significant and failed to detect the relevant suktas. SBERT detected four of the famous suktas as separate groups, but mistakenly combined three of them into a single mixed group. Also, the SBERT network was not statistically significant.

古籍分析语义网络文本挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。