为印度教至上学派经典打造可离线使用的逐词解析阅读器
A Word-Level Digital Reader of the Prasthanatrayi with Sankara's Bhasya: Corpus, Method, and an Open, Offline Reading Aid for the Advaita Vedanta Canon
- 基于规则与大模型结合的分词分析系统,精准拆解梵文音变与复合词
- 覆盖36,881个原文词、95,587个注疏词形,支持全文本词频检索
- 完全离线运行的HTML文件,适合研究者与学习者长期使用
《普拉斯坦塔里》——十部主要奥义书、梵经与薄伽梵歌,配以商羯罗注释(跋沙亚)——是不二论吠檀多的核心典籍。连续音变(沙地)、长复合词(萨马萨)及密集学术文体使逐词阅读困难:词边界与语法含义均被遮蔽。本文构建了一个开源、完全离线的《普拉斯坦塔里》及其商羯罗注疏的逐词数字阅读器。所有词(原文与注疏)均可点击,弹出框显示分词(帕达彻达)、形态分析与释义。每词含词根,使阅读器兼具词频索引功能:对词头搜索可查出其所有变体及复合词中的出现,跨越双层文本。资源涵盖13个注疏单元(2,971节经文、格言与散文段落;36,881个原文词分析),全局词表含95,587个注疏表面词形。我们描述了语料库、混合处理流程——基于规则的音变拆解+词形词典+实证语料查证,辅以大模型分析与对抗性双轮验证协议——及可持续的人工校验循环,校正结果在每次重建中保留。内在评估显示,高置信度分析与权威词形词典在99%以上实证形式上一致;盲评确认质量随置信度下降而系统性退化,错误集中于低置信度层级,恰为人工校验所聚焦。阅读器为单一自包含HTML文件,无需服务器或网络,免费分发,作为教学与阅读工具。
原文摘要 · Abstract (English)
The Prasthanatrayi -- the ten principal Upanisads, the Brahmasutra, and the Bhagavadgita, with Sankara's commentaries (bhasya) -- is the foundational corpus of Advaita Vedanta. Continuous euphonic combination (sandhi), long compounds (samasa), and dense scholastic prose make it hard to read at the word level: where one word ends, and what each word means grammatically, are both obscured. We present an open, fully offline, word-level digital reader of the entire Prasthanatrayi with Sankara's bhasya. Every word -- of both the root text (mula) and the commentary -- is clickable and resolves to a pop-up giving its split (padaccheda), morphological analysis, and gloss. Because every word carries a lemma, the reader also acts as a concordance: a search on a dictionary headword retrieves all of that word's inflected and sandhi-hidden occurrences, and its occurrences inside compounds, across both layers. The resource covers thirteen commentarial units (2,971 verses, sutras, and prose sections; 36,881 analysed word-occurrences of root text) and a global dictionary of 95,587 distinct commentarial surface forms. We describe the corpus, the hybrid pipeline -- a rule-based sandhi splitter over an inflected-form lexicon and attested-corpus look-ups, with LLM-assisted analysis under an adversarial two-pass verification protocol -- and a durable human-review loop whose corrections survive every regeneration. An intrinsic evaluation against independent Sanskrit resources finds high-confidence analyses agree with an authoritative inflectional lexicon on over 99% of attested forms, and a band-blind adjudication confirms that quality degrades predictably across confidence bands, with errors concentrated in the low-confidence tier the review loop targets. The reader is a single self-contained HTML file needing no server or network, offered as a freely redistributable teaching and reading aid.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。