arXiv:2606.08770cs.CLcs.AI2026-06中稿 · the 2nd Workshop o…

用Transformer和集成学习分析尼泊尔语梗图文本,提升仇恨言论与情感识别效果。

TeamHerald@CHIPSAL 2026: Hate Speech Detection and Sentiment Analysis of Nepali Memes using Transformer-based Architectures and Ensemble Learning

  • 通过OCR提取梗图文字,用Transformer模型建模文本特征。
  • 软投票集成在多分类任务上比单模型提升15.8%的宏平均F1分数。
  • 发现二分类与多分类任务应适配不同集成策略,指导实际应用选择。

尼泊尔语网络梗图分析因频繁的语言混杂及缺乏基准资源而困难重重。尽管梗图兼具视觉与文本元素,本研究采用以文本为中心的方法,通过OCR层提取嵌入文本,并使用基于Transformer的架构进行建模。我们评估了六种不同模型,并在两个任务上比较硬投票与软投票集成策略的性能:二分类仇恨言论检测与三分类情感分析。实验结果表明,单一解码器模型在二分类任务中表现最佳,而软投票集成在多分类任务中表现最优,相比最强单模型基线,宏平均F1得分提升15.8%。这些发现表明,集成策略在二分类与多分类任务中的表现存在差异,强调了根据分类目标选择合适聚合方法的重要性。

原文摘要 · Abstract (English)

The analysis of internet memes in the Nepali language is complicated by frequent code-mixing and a lack of established baseline resources. While memes inherently combine visual and textual elements, this study focuses on a text-centric approach by extracting embedded text using an OCR layer and modeling it with Transformer-based architectures. We evaluate six distinct models and investigate the comparative effectiveness of Hard and Soft Voting ensemble strategies across two tasks: binary hate speech detection and three-class sentiment analysis. Experimental results show that a standalone decoder-only model achieved the highest performance for binary classification, whereas the Soft Voting ensemble performed best for the multi-class sentiment task, yielding a 15.8% relative improvement in Macro F1-score over the strongest standalone baseline. These findings suggest that ensemble strategies behave differently across binary and multi-class tasks, highlighting the importance of selecting aggregation methods suited to the classification objective.

情感分析仇恨言论检测Transformer集成学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。