利用语言差异提升低资源方言的泛化能力,效果显著优于传统方法。
Harnessing Linguistic Dissimilarity for Language Generalization on Unseen Low-Resource Varieties

- 分两阶段框架:先选优质源语言,再用双分支网络分离特异性与共性特征
- 在10种低资源方言上依赖解析任务指标提升54.62%,验证跨品种泛化能力
- 适合研究多语言模型泛化、低资源语言处理的学者和工程师
在多语言大模型的发展中,特定群体使用的低资源语言变体仍被忽视。多数跨语言研究致力于对齐相似变体并缩小差异,但对于低资源变体而言,语言差异本身也是泛化至未见变体的重要线索。不同于以往方法,本文提出一种两阶段语言泛化框架:首先设计针对低资源变体的TOPPing源选择方法;其次提出轻量级VACAI-Bowl架构,通过双分支结构分别学习变体特异性属性与变体不变属性,后者采用对抗训练实现。我们在结构预测任务上评估该框架,此类任务可作为其他下游任务的代理指标。实验表明,结合TOPPing与VACAI-Bowl在依赖解析任务上平均提升54.62%,覆盖10种低资源变体。
原文摘要 · Abstract (English)
Low-resource language varieties used by specific groups remain neglected in the development of Multilingual Language Models. A great deal of cross-lingual research focuses on inter-lingual language transfer which strives to align allied varieties and minimize differences between them. However, for low-resource varieties, linguistic dissimilarity is also an important cue allowing generalization to unseen varieties. Unlike prior approaches, we propose a two-stage Language Generalization framework that focuses on capturing variety-specific cues while also exploiting rich overlap offered by high-resource source variety. First, we propose TOPPing, a source-selection method specifically designed for low-resource varieties. Second, we suggest a lightweight VACAI-Bowl architecture that learns variety-specific attributes with one branch while a parallel branch captures variety-invariant attributes using adversarial training. We evaluate our framework on structural prediction tasks, which are among the few tasks available, as proxy for performance on other downstream tasks. Using VACAI-Bowl with TOPPing yields an average 54.62% improvement in the dependency parsing task, which serves as a proxy for performance on other downstream tasks across 10 low-resource varieties.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。