为非洲语言社交媒体文本优化预训练模型,提升情感与仇恨言论识别效果
AfroXLMR-Social: Adapting Pre-trained Language Models for African Languages Social Media Text
- 采用领域和任务自适应预训练,提升低资源非洲语言模型性能
- 在19种语言上情感分析等任务的F1分数提升1%至30%
- 适用于非洲语言NLP研究者及多语言社会计算应用开发者
当前自然语言处理的进步依赖于多种来源构建的语言模型。然而,许多低资源语言的语料库多样性有限,且偏向宗教领域,导致在社交网络等快速变化的远域任务上表现不佳。领域自适应预训练(DAPT)和任务自适应预训练(TAPT)是通过持续预训练减少偏见的常用方法,但尚未应用于非洲多语言编码器。本文探索了DAPT与TAPT在非洲语言社交媒体领域的应用,提出AfriSocial——一个大规模社交媒体与新闻语料库,用于多语言持续预训练。基于此,DAPT在三个主观任务(情感分析、多标签情绪分类、仇恨言论识别)中将F1分数提升1%至30%,覆盖19种语言。TAPT在单任务数据上训练后,可提升其他相关任务表现,例如用未标注情感数据训练细粒度情绪分类任务,使基准结果提升0.55%至15.11%。结合两种方法(DAPT + TAPT)进一步提升整体性能。数据与模型资源已发布于HuggingFace。
原文摘要 · Abstract (English)
Language models built from various sources are the foundation of today's NLP progress. However, for many low-resource languages, the diversity of domains is often limited, more biased to a religious domain, which impacts their performance when evaluated on distant and rapidly evolving domains such as social media. Domain adaptive pre-training (DAPT) and task-adaptive pre-training (TAPT) are popular techniques to reduce this bias through continual pre-training for BERT-based models, but they have not been explored for African multilingual encoders. In this paper, we explore DAPT and TAPT continual pre-training approaches for African languages social media domain. We introduce AfriSocial, a large-scale social media and news domain corpus for continual pre-training on several African languages. Leveraging AfriSocial, we show that DAPT consistently improves performance (from 1% to 30% F1 score) on three subjective tasks: sentiment analysis, multi-label emotion, and hate speech classification, covering 19 languages. Similarly, leveraging TAPT on the data from one task enhances performance on other related tasks. For example, training with unlabeled sentiment data (source) for a fine-grained emotion classification task (target) improves the baseline results by an F1 score ranging from 0.55% to 15.11%. Combining these two methods (i.e. DAPT + TAPT) further improves the overall performance. The data and model resources are available at HuggingFace.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。