arXiv:2510.23585cs.CLcs.AI2025-10

比较传统模型与Transformer在社交媒体希望言论识别中的表现。

Hope Speech Detection in Social Media English Corpora: Performance of Traditional and Transformer Models

  • 用SVM、逻辑回归等传统模型和微调的Transformer对比
  • Transformer最佳模型宏F1达0.79,优于传统模型的0.78
  • 适合关注情感分析与小数据下语义捕捉的研究者

希望言论的识别已成为一项有前景的自然语言处理任务,旨在检测社交媒体平台上体现能动性与目标导向行为的激励性表达。本文评估了传统机器学习模型与微调的Transformer在先前划分的希望言论数据集上的表现,该数据集包含训练、开发和测试集。在开发集测试中,线性核SVM和逻辑回归均达到0.78的宏F1,RBF核SVM为0.77,朴素贝叶斯为0.75。而Transformer模型表现更优,最佳模型实现加权精确率0.82、加权召回率0.80、加权F1 0.79、宏F1 0.79,准确率0.80。结果表明,尽管配置良好的传统模型仍具高效性,但基于Transformer的架构能更好捕捉希望言论的细微语义,从而在精确率和召回率上取得更高表现,提示大型Transformer和大语言模型在小数据集上可能具备更强潜力。

原文摘要 · Abstract (English)

The identification of hope speech has become a promised NLP task, considering the need to detect motivational expressions of agency and goal-directed behaviour on social media platforms. This proposal evaluates traditional machine learning models and fine-tuned transformers for a previously split hope speech dataset as train, development and test set. On development test, a linear-kernel SVM and logistic regression both reached a macro-F1 of 0.78; SVM with RBF kernel reached 0.77, and Naïve Bayes hit 0.75. Transformer models delivered better results, the best model achieved weighted precision of 0.82, weighted recall of 0.80, weighted F1 of 0.79, macro F1 of 0.79, and 0.80 accuracy. These results suggest that while optimally configured traditional machine learning models remain agile, transformer architectures detect some subtle semantics of hope to achieve higher precision and recall in hope speech detection, suggesting that larges transformers and LLMs could perform better in small datasets.

希望言论情感分析Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。