为濒危的土耳其波马克语构建依赖句法解析资源并验证跨方言迁移效果。
Transfer Learning for an Endangered Slavic Variety: Dependency Parsing in Pomak Across Contact-Shaped Dialects
- 基于希腊方言数据零样本迁移至土耳其方言,评估音系与形态差异影响。
- 新标注650句土耳其方言语料,微调后准确率显著提升。
- 融合双方言数据的迁移学习策略,进一步优化解析性能。
本文针对濒危的东南斯拉夫语波马克语(主要在土耳其乌宗科普鲁地区使用)构建新的依赖句法解析资源与基线模型。该语言存在显著方言差异且无统一标准形式。研究聚焦于从希腊方言通用树库数据训练的解析器向土耳其方言的零样本迁移效果,量化音系与形态句法差异的影响。第二阶段引入一个包含650句的新手动标注土耳其方言语料,结果表明尽管规模小,但针对性微调可大幅提升准确率;进一步采用跨方言迁移学习结合两种方言数据,性能持续优化。实验揭示了方言间迁移的有效性及小样本语料的潜力。
原文摘要 · Abstract (English)
This paper presents new resources and baselines for Dependency Parsing in Pomak, an endangered Eastern South Slavic language with substantial dialectal variation and no widely adopted standard. We focus on the variety spoken in Turkey (Uzunköprü) and ask how well a dependency parser trained on the existing Pomak Universal Dependencies treebank, which was built primarily from the variety that is spoken in Greece, transfers across dialects. We run two experimental phases. First, we train a parser on the Greek-variety UD data and evaluate zero-shot transfer to Turkish-variety Pomak, quantifying the impact of phonological and morphosyntactic differences. Second, we introduce a new manually annotated Turkish-variety Pomak corpus of 650 sentences and show that, despite its small size, targeted fine-tuning substantially improves accuracy; performance is further boosted by cross-variety transfer learning that combines the two dialects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。