arXiv:2606.12649cs.CL2026-06

针对阿拉伯语心理疾病检测难题,提出自适应预训练与分层微调框架。

MentalMARBERT: Domain-Adaptive Pre-training and Two-Stage Fine-Tuning for Arabic Mental Health Disorders Detection

论文配图:MentalMARBERT: Domain-Adaptive Pre-training and Two-Stage Fine-Tuning for Arabic Mental Health Disorders Detection
图 1 · 摘自论文原文
  • 用未标注的阿拉伯心理健康推文进行领域自适应预训练
  • 分层两阶段微调使宏F1达0.861,准确率0.877
  • 适用于阿拉伯语心理健康文本分析的研究者与应用开发者

由于方言差异、非正式语言、高质量标注资源稀缺及类别严重不平衡,从阿拉伯语社交媒体文本中检测心理疾病仍具挑战。尽管英语在心理健康自然语言处理方面进展显著,阿拉伯语多类别障碍分类仍研究不足。本研究提出一种两阶段框架:第一阶段对AraBERT、CAMeLBERT和MARBERT三个预训练模型使用大规模未标注阿拉伯心理健康推文进行领域自适应与任务自适应预训练(DAPT和TAPT),并统一评估以选出最优骨干模型;第二阶段在四种配置下评估所选模型,包括单阶段与分层两阶段分类架构,结合全微调与低秩适配(LoRA)。为支持本研究,构建了一个包含50,670条推文、六类标注的新阿拉伯语心理健康数据集,标注者间一致性高(Krippendorff's Alpha = 0.733,平均成对一致率为0.797)。实验结果表明,经领域自适应的MARBERT(MentalMARBERT)在准确率与宏F1上均显著优于基线模型。分层两阶段架构结合全微调表现最佳,宏F1达0.861,准确率为0.877。结果验证了领域特定自适应预训练与分层分类的有效性。

原文摘要 · Abstract (English)

Detecting mental health disorders from Arabic social media text remains challenging due to dialectal variation, informal language, limited high-quality annotated resources, and severe class imbalance. While English mental health natural language processing (NLP) has progressed substantially, Arabic multi-class disorder classification remains insufficiently studied. This study proposes a two-phase framework for Arabic mental health text classification. In phase 1, three Arabic pre-trained language models, AraBERT, CAMeLBERT, and MARBERT, undergo Domain-Adaptive and Task-Adaptive Pretraining (DAPT and TAPT) using a large-scale corpus of unlabeled Arabic mental health tweets. The adapted models are evaluated under a unified protocol to identify the most effective backbone model. In phase 2, the selected model is assessed across four configurations combining single-stage and hierarchical two-stage classification architectures with full fine-tuning and Low-Rank Adaptation (LoRA). To support this study, we constructed a novel annotated Arabic mental health dataset comprising 50,670 tweets across six categories, with strong inter annotator agreement (Krippendorff's Alpha = 0.733, average pairwise agreement = 0.797). Experimental results show that the domain-adapted MARBERT (MentalMARBERT) achieves statistically significant improvements over baseline models in both accuracy and macro-F1. The hierarchical two-stage architecture combined with full fine-tuning achieves the best overall performance, reaching a macro-F1 of 0.861 and an accuracy of 0.877. These findings demonstrate the effectiveness of domain-specific adaptive pretraining and hierarchical classification for Arabic mental health disorder detection.

心理检测阿拉伯语NLP预训练模型分层分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。