用深度学习分析拉丁语到奥克语的语法性别演变机制
Lost in Translation? Exploring the Shift in Grammatical Gender from Latin to Occitan

- 设计可解释的深度学习框架,融合词法与语境特征
- 发现词干信息对性别预测贡献显著高于上下文
- 适合历史语言学与计算语言学研究者参考
从拉丁语到罗曼语族的语言历时演变中,语法性别系统从三性(阳性、阴性、中性)重组为多数罗曼语中的两性(阳性、阴性)。本文提出一种可解释的深度学习框架,从词法与语境两个层面探究该现象。首先,证明传统分词策略在低资源历史语料中表现不足,所提分词器性能更优。在词法层面,评估形态特征对性别预测的贡献;在语境层面,量化不同词性类别对性别预测的影响。综合分析揭示了性别信息在词干与其句法上下文间的分布规律。代码、数据集及结果已公开于 https://github.com/ahan-2000/Lost-in-Translation-。
原文摘要 · Abstract (English)
The diachronic evolution from Latin to the Romance languages involved a restructuring of the grammatical gender system from a tripartite configuration (masculine, feminine, neuter) to a bipartite one (masculine, feminine) in most Romance languages. In this work, we introduce an interpretable deep learning framework to investigate this phenomenon at both lexical and contextual levels. First, we show that conventional tokenization strategies are insufficiently robust for this low-resource historical setting, and that our proposed tokenizer improves performance over these baselines. At the lexical level, we evaluate the contribution of morphological features to gender prediction. At the contextual level, we quantify the contributions of different part-of-speech categories to grammatical gender prediction. Together, these analyses characterize the distribution of gender information between the lemma and its sentential context. We make our codebase, datasets, and results publicly available at \href{https://github.com/ahan-2000/Lost-in-Translation-}{https://github.com/ahan-2000/Lost-in-Translation-}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。