用梵语语法框架统一印地语系语言处理,提升模型效率与可迁移性。
A Pāninian Foundation for Indic Language Processing
- 基于两千年的梵语语法体系构建统一语言架构
- 提出四部分基准测试套件,实现跨语言资源融合
- 适合研究印地语系语言、神经模型可解释性的学者
超过十亿人使用印地语系语言,但其自然语言处理基础设施仍零散且发展不足。根源在于当前方法按语言或语系小分支分别构建分析器、解析器和数据集,重复建设。而事实上,历经两千年围绕梵语的演化,印地语系语言共享了由婆罗摩笈多的《八章书》形式化的形态句法架构。该结构跨越语系界限,形成统一框架。我们主张这一婆罗摩笈多框架可作为领域缺失的统一计算架构,以之为基础的基准测试将使系统更准确、更高效、更具迁移能力,有效将多个稀疏的印地语系资源合并为单一高资源通用语言基础。为此,我们提出一个四部分基准测试套件,使这一共享架构显式化、可度量并可用于实际应用。此外,我们提出关键问题:在这些语言上训练的神经模型是否自发地表征了婆罗摩笈多的语法范畴?
原文摘要 · Abstract (English)
More than a billion people communicate in Indic languages, yet the natural language processing infrastructure serving them remains fragmented and underdeveloped. The cause is structural: the field organizes its tools and benchmarks around individual languages or small subsets of genealogical language families, building separate analyzers, parsers, and datasets for each language and starting over for the next. This overlooks a deep regularity. Through more than two millennia of convergence around Sanskrit, Indic languages came to share a morphosyntactic architecture formalized in Pānini's grammar, the Astādhyāyī. This cuts across genealogical lines, uniting languages through a common framework. We argue that this Pāninian framework supplies a unifying computational architecture the field has lacked, and that benchmarks grounded explicitly in it would make Indic language systems more accurate, more data-efficient, and more transferable, effectively merging many apparently disparate and sparse Indic language resources into a single high-resource metalanguage bedrock. We propose a four-part benchmark suite to render this shared architecture explicit, measurable, and ready to be leveraged for practical applications. Moreover, we underscore the question it raises for interpretability research: whether neural models trained on these languages come to represent Pānini's categories on their own.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。