用软提示提升小模型跨语言零样本分类能力
Enhancing Small Language Models for Cross-Lingual Generalized Zero-Shot Classification with Soft Prompt Tuning
- 设计轻量软提示,利用高资源语言提升低资源语言分类
- 在106种语言上实现强跨语言迁移与未见类别泛化
- 适合资源匮乏语言的零样本分类任务
在自然语言处理中,零样本分类(ZSC)使模型能在无训练标签的情况下对未见类别进行分类,尤其适用于低资源语言和领域。尽管预训练语言模型(PLMs)在ZSC中表现良好,但通常依赖大规模训练数据或外部知识,限制了其在多语言和低资源场景的应用。近期基于自然语言提示的方法减少了对大量训练数据的依赖,但在整合相关任务的标注数据时效果不佳,尤其当数据来自不同语言或分布时。此外,现有提示方法多依赖特定语言的手动构造提示,难以适应跨语言设置。为此,我们提出RoSPrompt,一种轻量且数据高效的软提示训练方法,增强小规模多语言PLM在跨语言零样本分类中的表现,并确保对数据分布变化的鲁棒泛化能力。RoSPrompt专为小规模多语言PLM设计,使其能利用高资源语言提升低资源语言性能,无需大量微调或高计算成本。我们在涵盖106种语言的多个多语言PLM和数据集上评估该方法,验证了其强大的跨语言迁移性能和对未见类别的鲁棒泛化能力。
原文摘要 · Abstract (English)
In NLP, Zero-Shot Classification (ZSC) has become essential for enabling models to classify text into categories unseen during training, particularly in low-resource languages and domains where labeled data is scarce. While pretrained language models (PLMs) have shown promise in ZSC, they often rely on large training datasets or external knowledge, limiting their applicability in multilingual and low-resource scenarios. Recent approaches leveraging natural language prompts reduce the dependence on large training datasets but struggle to effectively incorporate available labeled data from related classification tasks, especially when these datasets originate from different languages or distributions. Moreover, existing prompt-based methods typically rely on manually crafted prompts in a specific language, limiting their adaptability and effectiveness in cross-lingual settings. To address these challenges, we introduce RoSPrompt, a lightweight and data-efficient approach for training soft prompts that enhance cross-lingual ZSC while ensuring robust generalization across data distribution shifts. RoSPrompt is designed for small multilingual PLMs, enabling them to leverage high-resource languages to improve performance in low-resource settings without requiring extensive fine-tuning or high computational costs. We evaluate our approach on multiple multilingual PLMs across datasets covering 106 languages, demonstrating strong cross-lingual transfer performance and robust generalization capabilities over unseen classes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。