为低资源古荷兰语构建基于Transformer的句法解析器,提升跨领域泛化能力。
Lexicalized Constituency Parsing for Middle Dutch: Low-resource Training and Cross-Domain Generalization
- 用Transformer模型适配古荷兰语,结合高资源语言联合训练
- 跨域性能提升0.73 F1,地理时间接近的语言增益最大
- 200例/领域是有效跨域训练的最低阈值,优于传统PCFG
近年来,神经网络与上下文词嵌入在历史语言句法分析中受到关注,但多数进展集中于依存解析,古荷兰语等低资源历史语言的成分解析研究较少。本文将基于Transformer的成分解析器应用于高度异质且数据稀少的古荷兰语,探索提升其领域内与跨领域性能的方法。实验表明,与高资源辅助语言联合训练可使F1得分最高提升0.73,地理和时间上更接近古荷兰语的语言带来最大增益。进一步评估新标注多领域数据的利用策略,发现微调与数据合并效果相当,且神经解析器始终优于当前使用的基于PCFG的解析器。此外,通过特征分离技术探索领域适应,证明每领域至少需约200个样本才能有效提升跨域表现。
原文摘要 · Abstract (English)
Recent years have seen growing interest in applying neural networks and contextualized word embeddings to the parsing of historical languages. However, most advances have focused on dependency parsing, while constituency parsing for low-resource historical languages like Middle Dutch has received little attention. In this paper, we adapt a transformer-based constituency parser to Middle Dutch, a highly heterogeneous and low-resource language, and investigate methods to improve both its in-domain and cross-domain performance. We show that joint training with higher-resource auxiliary languages increases F1 scores by up to 0.73, with the greatest gains achieved from languages that are geographically and temporally closer to Middle Dutch. We further evaluate strategies for leveraging newly annotated data from additional domains, finding that fine-tuning and data combination yield comparable improvements, and our neural parser consistently outperforms the currently used PCFG-based parser for Middle Dutch. We further explore feature-separation techniques for domain adaptation and demonstrate that a minimum threshold of approximately 200 examples per domain is needed to effectively enhance cross-domain performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。