arXiv:2606.18222cs.CLcs.DL2026-06

构建跨教派哲学评论语料库,支持多视角文本对比分析。

Darshana Graph: A Parallel Commentary Corpus for Comparative Indian Philosophy, with Stylometric and Exploratory Graph Analyses

  • 收录12.5万条经典文本,8500条经文配十八位注释家对照
  • 发现引文密度与反驳率呈中度负相关,传承脉络中反驳率上升明显
  • 用规则引擎提取哲学关系图,适合比较哲学与数字人文研究者

我们提出Darshana Graph,一个包含超过12.5万条记录的跨宗教哲学语料库,涵盖印度教、佛教与耆那教的经典传统,来源包括《薄伽梵歌》《梵经》《奥义书》《巴利圣典》及核心耆那教文献。其独特之处在于约8500条梵文经文与十八位历史注释家(代表五大学派与其它学说)的对应记录,实现对同一原文的跨注释者直接比对。据我们所知,目前无公开资源提供如此规模的跨注释者对齐。我们基于该语料库开展两项分析:一是无需机器学习的风格计量分析,通过引用密度、明确反驳率和句子复杂度衡量论证风格,发现引用密度与反驳率存在中度负相关,同源学派中三位注释者的反驳率显著升高,且巴利圣典内部存在可测的文体差异;二是使用预定义关系词汇表与确定性后处理验证的受限大模型流水线,提取概念间类型化哲学关系,生成的关系图揭示了跨学派分歧模式,但也暴露了嵌入式分析与图推导结果不一致的问题。全部语料、关系图与源代码均已开放。

原文摘要 · Abstract (English)

We introduce Darshana Graph, a corpus of over 125,000 text records spanning classical Hindu, Buddhist, and Jain philosophical traditions, drawn from public-domain and openly licensed translations of sources including the Bhagavad Gita, Brahma Sutras, principal Upanishads, the Pali Canon, and core Jain texts. Its distinctive contribution lies in a structurally unique subset of roughly 8,500 Hindu and Jain records in which the same root verse or sutra is aligned across eighteen historical commentators representing five schools of Vedanta and other darshanas, enabling direct comparison of how independent interpretive traditions read identical source material. To our knowledge, no publicly available resource provides comparable cross-commentator alignment at this scale. We present two analyses built on this corpus. First, a transparent stylometric comparison requiring no machine learning measures argumentative style through scriptural citation density, explicit refutation rate, and sentence complexity. It finds a moderate negative correlation between citation density and refutation rate, a marked increase in refutation rate across three commentators in a related doctrinal lineage, and measurable genre-level differences within the Pali Canon itself. Second, we describe a constrained large language model pipeline that extracts typed philosophical relationships between concepts using a predefined relation vocabulary and deterministic post-hoc validation. The resulting graph surfaces cross-school disagreement patterns while also revealing important extraction limitations, including cases where an independent embedding-based analysis disagrees with the graph-derived findings. We release the full corpus, extracted relationship graph, and all source code.

哲学语料跨学派比较文本分析数字人文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。