arXiv:2508.19464cs.CLcs.AI2025-08

用对比学习提升低资源语言的少样本迁移能力

Bridging Language Gaps: Enhancing Few-Shot Language Adaptation

  • 通过对比学习对齐跨语言表征,实现知识高效迁移
  • 在少样本条件下显著超越现有跨语言迁移方法
  • 适合需要快速适配新语言的多语言NLP应用

多语言自然语言处理面临语言资源不均衡问题:高资源语言拥有丰富数据,而低资源语言缺乏足够标注数据进行有效训练。本文提出对比学习与提示结合的方法(CoLAP),通过融合对比学习与跨语言表征,实现从高资源语言到低资源语言的任务特定知识迁移。该方法具有数据高效性,可在极少标注数据下快速适应新语言,减少对大规模标注集的依赖。我们在包含自然语言推理和关系抽取在内的多项自然语言理解任务上,对仅编码器和仅解码器的多语言模型进行了实验,覆盖高低资源语言。结果表明,CoLAP在少样本跨语言迁移场景下优于现有基线方法及上下文学习,显著缩小了跨语言性能差距,为更高效的多语言NLP技术发展提供了支持。

原文摘要 · Abstract (English)

The disparity in language resources poses a challenge in multilingual NLP, with high-resource languages benefiting from extensive data, while low-resource languages lack sufficient data for effective training. Our Contrastive Language Alignment with Prompting (CoLAP) method addresses this gap by integrating contrastive learning with cross-lingual representations, facilitating task-specific knowledge transfer from high-resource to lower-resource languages. The primary advantage of our approach is its data efficiency, enabling rapid adaptation to new languages and reducing the need for large labeled datasets. We conduct experiments with multilingual encoder-only and decoder-only language models on natural language understanding tasks, including natural language inference and relation extraction, evaluating performance across both high- and low-resource languages. Our results demonstrate that CoLAP outperforms few-shot cross-lingual transfer baselines and in-context learning, even with limited available data. This effectively narrows the cross-lingual performance gap, contributing to the development of more efficient multilingual NLP techniques.

少样本学习跨语言迁移对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。