用因果方法提升低资源语言模型在分布偏移下的泛化能力
When Distributions Shifts: Causal Generalization for Low-Resource Languages
- 通过GPT-4o-mini生成反事实改写句增强数据
- 在17种语言上实现跨域稳定性能提升
- 适合低资源语言与鲁棒性要求高的场景
机器学习模型在分布偏移下表现不佳,这一问题在低资源语境中尤为突出。本文研究两种因果领域泛化方法在低资源自然语言处理中的应用。首先,利用GPT-4o-mini生成反事实改写句,对约鲁巴语和伊博语的NaijaSenti Twitter语料进行情感分类的数据增强。其次,将去偏方面评论(DINER)框架扩展至多语言场景,构建了由SemEval-2014任务翻译而来的17种语言的Afri-SemEval数据集,并采用因果不变表示学习。实验表明,反事实增强带来一致性能提升,因果表示学习显著改善跨域泛化能力,在多种语言上均有效。
原文摘要 · Abstract (English)
Machine learning models often fail under distribution shifts, a problem exacerbated in low-resource settings where limited data restricts robust generalization. Domain generalization(DG) methods address this challenge by learning representations that remain invariant across domains, frequently leveraging causal principles. In this work, we study two causal DG approaches for low-resource natural language processing. First, we apply causal data augmentation using GPT-4o-mini to generate counterfactual paraphrases for sentiment classification on the NaijaSenti Twitter corpus in Yoruba and Igbo. Second, we investigate invariant causal representation learning with the Debiasing in Aspect Review (DINER) framework for aspect-based sentiment analysis. We extend DINER to a multilingual setting by introducing Afri-SemEval, a dataset of 17 languages translated from SemEval-2014 Task. Experiments show improved robustness to unseen domains, with consistent gains from counterfactual augmentation and enhanced out-of-distribution performance from causal representation learning across multiple languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。