arXiv:2509.25752cs.CLcs.LG2025-09

多语言希望话语分类,提升社交平台正向内容识别

Detecting Hope Across Languages: Multiclass Classification for Positive Online Discourse

  • 用XLM-RoBERTa模型实现英、乌尔都、西语希望话语三类分类
  • 在PolyHope数据集上达到领先宏F1分数,尤其在低资源语言表现优
  • 适合做跨语言正向言论监测的系统开发者与社会计算研究者

社交媒体中希望话语的检测已成为促进积极对话与心理福祉的关键任务。本文提出一种基于机器学习的多语言多类别希望话语检测方法,涵盖英语、乌尔都语和西班牙语。利用XLM-RoBERTa等变压器模型,将希望话语分为三类:普遍性希望、现实性希望与非现实性希望。该方法在PolyHope-M 2025共享任务的PolyHope数据集上进行评估,整体表现优异。与现有模型对比显示,本方法在宏平均F1得分上显著超越先前最先进水平。同时讨论了低资源语言中检测的挑战及泛化能力提升潜力。本工作推动了多语言细粒度希望话语检测模型的发展,可应用于正向内容审核与支持性在线社区建设。

原文摘要 · Abstract (English)

The detection of hopeful speech in social media has emerged as a critical task for promoting positive discourse and well-being. In this paper, we present a machine learning approach to multiclass hope speech detection across multiple languages, including English, Urdu, and Spanish. We leverage transformer-based models, specifically XLM-RoBERTa, to detect and categorize hope speech into three distinct classes: Generalized Hope, Realistic Hope, and Unrealistic Hope. Our proposed methodology is evaluated on the PolyHope dataset for the PolyHope-M 2025 shared task, achieving competitive performance across all languages. We compare our results with existing models, demonstrating that our approach significantly outperforms prior state-of-the-art techniques in terms of macro F1 scores. We also discuss the challenges in detecting hope speech in low-resource languages and the potential for improving generalization. This work contributes to the development of multilingual, fine-grained hope speech detection models, which can be applied to enhance positive content moderation and foster supportive online communities.

希望话语多语言情感分析内容审核

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。