arXiv:2509.20315cs.CLcs.LG2025-09

用主动学习提升多语言希望言论检测效果

Multilingual Hope Speech Detection: A Comparative Study of Logistic Regression, mBERT, and XLM-RoBERTa with Active Learning

  • 结合主动学习与多语言Transformer模型
  • XLM-RoBERTa在多语言数据上准确率最高
  • 小样本下仍保持良好性能,适合资源少场景

希望言论通过鼓励和乐观促进网络积极对话,但在多语言和低资源环境下检测仍具挑战。本文提出一种基于主动学习的多语言希望言论检测框架,采用mBERT和XLM-RoBERTa等Transformer模型,在英语、西班牙语、德语和乌尔都语数据集上进行实验,涵盖近期共享任务的基准测试集。结果表明,Transformer模型显著优于传统基线,其中XLM-RoBERTa整体准确率最高。此外,主动学习策略在小规模标注数据下仍保持优异性能。本研究验证了多语言Transformer与数据高效训练策略结合的有效性。

原文摘要 · Abstract (English)

Hope speech language that fosters encouragement and optimism plays a vital role in promoting positive discourse online. However, its detection remains challenging, especially in multilingual and low-resource settings. This paper presents a multilingual framework for hope speech detection using an active learning approach and transformer-based models, including mBERT and XLM-RoBERTa. Experiments were conducted on datasets in English, Spanish, German, and Urdu, including benchmark test sets from recent shared tasks. Our results show that transformer models significantly outperform traditional baselines, with XLM-RoBERTa achieving the highest overall accuracy. Furthermore, our active learning strategy maintained strong performance even with small annotated datasets. This study highlights the effectiveness of combining multilingual transformers with data-efficient training strategies for hope speech detection.

多语言希望言论主动学习Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。