用规则引导的进化搜索重建原始词形,比传统方法更准确。
Unsupervised Protoform Reconstruction through Parsimonious Rule-guided Heuristics and Evolutionary Search
- 结合统计规律与语言学规则,在进化框架中优化词形重建
- 在拉丁语原始词形重建任务中,字符准确率和发音合理性均显著提升
- 适合对历史语言学或自动词源重建感兴趣的读者
我们提出一种无监督方法,用于重建原始词形(protoforms),即现代语言词形的祖先形式。以往研究主要依赖音变的概率模型从同源词集中推断原始词形,但这类方法受限于数据驱动的特性。本文方法将数据驱动推理与基于规则的启发式策略相结合,构建于进化优化框架中,同时利用统计模式和语言学约束来指导重建过程。我们在五种罗曼语的同源词数据集上评估该方法,用于重建拉丁语原始词形。实验结果表明,该方法在字符级准确率和发音合理性指标上均显著优于现有基线。
原文摘要 · Abstract (English)
We propose an unsupervised method for the reconstruction of protoforms i.e., ancestral word forms from which modern language forms are derived. While prior work has primarily relied on probabilistic models of phonological edits to infer protoforms from cognate sets, such approaches are limited by their predominantly data-driven nature. In contrast, our model integrates data-driven inference with rule-based heuristics within an evolutionary optimization framework. This hybrid approach leverages on both statistical patterns and linguistically motivated constraints to guide the reconstruction process. We evaluate our method on the task of reconstructing Latin protoforms using a dataset of cognates from five Romance languages. Experimental results demonstrate substantial improvements over established baselines across both character-level accuracy and phonological plausibility metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。