用计算风格学分析佛教三藏英文译本,发现不同文本词汇分布差异显著。
Three Buddhist Vocabularies: Computational Stylometry of the English Pali Canon across Sutta, Vinaya, and Abhidhamma
- 通过词频分布、词汇多样性等指标对比三藏译文风格
- 《阿毗达摩概要》词汇最丰富,数字词密度最高
- 不同译本间词汇重合度低,反映翻译策略差异
我们对《巴利三藏》全部三藏的英文译本进行计算风格学分析,扩展了此前仅针对经藏的研究。语料库包含134,831个段落,涵盖苏贾托比丘的经藏(114,591段,CC0)、布拉马利比丘的律藏(7,923段,CC0 2026)、I.B. 霍纳1938年译本(2,826段)、《阿毗达摩概要》三个英译本(2,077段),以及法藏与说一切有部的跨传承律藏文本。计算了Zipf幂律指数(OLS拟合,R² > 0.989)、移动平均型词汇多样性(MATTR-500)、数字词密度及词汇重合度(Jaccard与Szymkiewicz-Simpson系数)。主要发现:(1) 所有语料均符合Zipf分布,律藏最接近理想斜率-1,而《概要》偏离最大,'意识'在第8位取代语法虚词;(2) MATTR-500显示经藏与律藏泰族译本词汇多样性相近(0.399和0.400),《概要》更高(0.560),经控量抽样验证;(3) 《概要》数字词密度达3.26%,与其系统枚举心物范畴一致;(4) 说一切有部律藏与泰族律藏词汇重合率为20.0%(Jaccard)和49.1%(重合系数),反映跨两千年法律传统共性;(5) 同一来源文本的两个英译本(相隔88年)仅共享24.2%词汇,如'沉思'与'入定'对应jhana,'失败'与'驱逐'对应parajika为关键差异。所有结果为点估计,未做显著性检验。代码与数据作为Darshana Graph语料库的开源扩展发布(arXiv:2606.18222)。
原文摘要 · Abstract (English)
We present a computational stylometric analysis of the Tipitaka across all three Pitakas in English translation, extending earlier work on the Sutta Pitaka alone. The corpus spans 134,831 segments from Bhikkhu Sujato's Sutta Pitaka (114,591 segments, CC0), Bhikkhu Brahmali's Vinaya Pitaka (7,923 segments, CC0 2026), I.B. Horner's 1938 Vinaya translation (2,826 segments), three English translations of the Abhidhammattha Sangaha compendium (2,077 segments), and cross-tradition Vinaya texts from the Dharmaguptaka and Mulasarvastivada schools. We compute Zipf rank-frequency distributions with OLS-fitted exponents, Moving Average TTR (MATTR-500), numeral-word density, and vocabulary overlap (Jaccard and Szymkiewicz-Simpson coefficients). Main findings: (1) all corpora show Zipf-consistent distributions (R2 > 0.989); the Vinaya is closest to ideal Zipf slope -1 and the Sangaha corpus deviates most, with 'consciousness' displacing grammatical particles at rank 8; (2) MATTR-500 shows the Sutta and Vinaya Theravada are nearly identical in lexical diversity (0.399 and 0.400), while the Sangaha corpus is genuinely more diverse (0.560), confirmed by size-controlled subsampling; (3) the Sangaha corpus has the highest numeral-word density (3.26%), consistent with its systematic enumeration of mental and material categories; (4) the Mulasarvastivada Vinaya shares 20.0% vocabulary (Jaccard) and 49.1% (overlap coefficient) with the Theravada Vinaya, reflecting shared legal heritage across two millennia; (5) two English translations of the same Vinaya source text share only 24.2% of their vocabulary across 88 years, with 'musing' versus 'absorption' for jhana and 'defeat' versus 'expulsion' for parajika as the most diagnostic shifts. All results are point estimates; no significance testing is conducted. Code and data are released as open-source extensions to the Darshana Graph corpus (arXiv:2606.18222).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。