arXiv:2512.05364cs.CL2025-12被引 1

用神经符号方法分析梵语2000年演变,发现复杂性转移而非简化。

Transformer-Enabled Diachronic Analysis of Vedic Sanskrit: Neural Methods for Quantifying Types of Language Change

  • 结合规则与深度学习生成伪标签,解决低资源语言数据不足问题。
  • 在147万词语料上实现52.4%特征检测率,准确识别复杂性转移。
  • 结果揭示语法复杂性向构词和哲学术语转移,适合计算语言学研究者。

本研究展示混合神经-符号方法如何为一种形态丰富、低资源的语言演化提供新见解。通过量化分析超过2000年的梵语演变,挑战了语言变化即简化的常识假设。我们利用100多个高精度正则表达式模式生成伪标签,微调多语言BERT模型,并通过新型置信度加权集成融合符号与神经输出,构建可扩展且可解释的系统。该框架应用于147万词的历时语料库,集成模型整体特征检测率达52.4%。研究发现,梵语整体形态复杂性并未下降,而是动态再分配:早期动词特征呈现周期性衰退,复杂性转移至其他领域,表现为构词大幅扩张及新哲学术语涌现。关键的是,系统生成校准良好的不确定性估计,置信度与准确率强相关(皮尔逊r = 0.92),总体校准误差低(ECE = 0.043),增强了计算文献学研究的可靠性。

原文摘要 · Abstract (English)

This study demonstrates how hybrid neural-symbolic methods can yield significant new insights into the evolution of a morphologically rich, low-resource language. We challenge the naive assumption that linguistic change is simplification by quantitatively analyzing over 2,000 years of Sanskrit, demonstrating how weakly-supervised hybrid methods can yield new insights into the evolution of morphologically rich, low-resource languages. Our approach addresses data scarcity through weak supervision, using 100+ high-precision regex patterns to generate pseudo-labels for fine-tuning a multilingual BERT. We then fuse symbolic and neural outputs via a novel confidence-weighted ensemble, creating a system that is both scalable and interpretable. Applying this framework to a 1.47-million-word diachronic corpus, our ensemble achieves a 52.4% overall feature detection rate. Our findings reveal that Sanskrit's overall morphological complexity does not decrease but is instead dynamically redistributed: while earlier verbal features show cyclical patterns of decline, complexity shifts to other domains, evidenced by a dramatic expansion in compounding and the emergence of new philosophical terminology. Critically, our system produces well-calibrated uncertainty estimates, with confidence strongly correlating with accuracy (Pearson r = 0.92) and low overall calibration error (ECE = 0.043), bolstering the reliability of these findings for computational philology.

语言演化神经符号梵语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。