arXiv:2512.22705cs.CLcs.AI2025-12中稿 · and presented at t…

用多语言模型检测低资源语种中的希望话语,提升网络正向交流。

GHaLIB: A Multilingual Framework for Hope Speech Detection in Low-Resource Languages

  • 基于XLM-RoBERTa等预训练模型构建多语言检测框架
  • 乌尔都语二分类F1达95.2%,多分类65.2%
  • 首次在乌尔都语等低资源语言中实现高效希望话语识别

希望话语在自然语言处理中仍属研究空白,现有工作主要集中在英语,导致乌尔都语等低资源语言缺乏相关工具。尽管基于Transformer的模型在仇恨和攻击性言论检测中表现良好,但其在希望话语检测中的应用仍有限,且跨语言验证不足。本文提出一个面向乌尔都语的多语言希望话语检测框架,采用XLM-RoBERTa、mBERT、EuroBERT及UrduBERT等预训练模型,结合简单预处理进行分类器训练。在PolyHope-M 2025基准测试中,该框架在乌尔都语二分类任务上达到95.2%的F1分数,多分类任务为65.2%,并在西班牙语、德语和英语中取得相当竞争力的结果。实验表明,现有多语言模型可在低资源环境下有效应用于希望话语识别,助力构建更积极的数字对话环境。

原文摘要 · Abstract (English)

Hope speech has been relatively underrepresented in Natural Language Processing (NLP). Current studies are largely focused on English, which has resulted in a lack of resources for low-resource languages such as Urdu. As a result, the creation of tools that facilitate positive online communication remains limited. Although transformer-based architectures have proven to be effective in detecting hate and offensive speech, little has been done to apply them to hope speech or, more generally, to test them across a variety of linguistic settings. This paper presents a multilingual framework for hope speech detection with a focus on Urdu. Using pretrained transformer models such as XLM-RoBERTa, mBERT, EuroBERT, and UrduBERT, we apply simple preprocessing and train classifiers for improved results. Evaluations on the PolyHope-M 2025 benchmark demonstrate strong performance, achieving F1-scores of 95.2% for Urdu binary classification and 65.2% for Urdu multi-class classification, with similarly competitive results in Spanish, German, and English. These results highlight the possibility of implementing existing multilingual models in low-resource environments, thus making it easier to identify hope speech and helping to build a more constructive digital discourse.

希望话语多语言低资源文本分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。