首个针对混合拼写乌尔都语的希望话语检测研究,填补低资源语言NLP空白
Hope Speech Detection in code-mixed Roman Urdu tweets: A Positive Turn in Natural Language Processing
- 构建首个混合拼写乌尔都语希望话语多分类数据集
- 自研注意力Transformer模型在交叉验证中达0.78准确率
- 适合关注低资源语言、社会情绪分析的研究者
希望是一种积极情绪,包含对未来有利结果的期待;希望话语指在不利情境下传递乐观、韧性与支持的表达。尽管希望话语检测在自然语言处理中受到关注,但现有研究主要集中在高资源语言和标准书写形式,忽视了非正式且资源匮乏的语言变体,如混合拼写乌尔都语。据我们所知,这是首个针对混合拼写乌尔都语中希望话语检测的研究,提出一个精心标注的数据集,填补了包容性NLP在低资源、非正式语言中的关键空白。本研究贡献包括:(1) 构建首个多类别标注数据集,涵盖泛化希望、现实希望、不切实际希望及非希望四类;(2) 探索希望的心理基础,分析其在混合拼写乌尔都语中的语言模式以指导数据构建;(3) 提出专为乌尔都语句法语义多样性优化的注意力机制Transformer模型,采用5折交叉验证评估;(4) 通过t检验验证性能提升的统计显著性。所提模型XLM-R表现最佳,交叉验证得分为0.78,优于基线SVM(0.75)和BiLSTM(0.76),分别提升4%和2.63%。
原文摘要 · Abstract (English)
Hope is a positive emotional state involving the expectation of favorable future outcomes, while hope speech refers to communication that promotes optimism, resilience, and support, particularly in adverse contexts. Although hope speech detection has gained attention in Natural Language Processing (NLP), existing research mainly focuses on high-resource languages and standardized scripts, often overlooking informal and underrepresented forms such as Roman Urdu. To the best of our knowledge, this is the first study to address hope speech detection in code-mixed Roman Urdu by introducing a carefully annotated dataset, thereby filling a critical gap in inclusive NLP research for low-resource, informal language varieties. This study makes four key contributions: (1) it introduces the first multi-class annotated dataset for Roman Urdu hope speech, comprising Generalized Hope, Realistic Hope, Unrealistic Hope, and Not Hope categories; (2) it explores the psychological foundations of hope and analyzes its linguistic patterns in code-mixed Roman Urdu to inform dataset development; (3) it proposes a custom attention-based transformer model optimized for the syntactic and semantic variability of Roman Urdu, evaluated using 5-fold cross-validation; and (4) it verifies the statistical significance of performance gains using a t-test. The proposed model, XLM-R, achieves the best performance with a cross-validation score of 0.78, outperforming the baseline SVM (0.75) and BiLSTM (0.76), with gains of 4% and 2.63% respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。