arXiv:2605.06231cs.CL2026-05ACL被引 1

用混合模型检测22种语言的网络极化,提升跨文化内容识别准确率。

YEZE at SemEval-2026 Task 9: Detecting Multilingual, Multicultural and Multievent Online Polarization via Heterogeneous Ensembling

论文配图:YEZE at SemEval-2026 Task 9: Detecting Multilingual, Multicultural and Multievent Online Polarization via Heterogeneous Ensembling
图 1 · 摘自论文原文
  • 融合XLM-RoBERTa和mDeBERTa构建异构集成模型
  • 独立建模+类别加权使极化检测效果更优
  • 适用于多语言、多文化、多事件场景的舆情分析

本文介绍我们在SemEval-2026任务9中的系统:检测多语言、多文化及多事件在线极化。该系统在22种语言中完成三项子任务:二元极化检测、目标分类与表现形式识别。提出一种由多语言预训练模型组成的异构集成方法,结合XLM-RoBERTa-large与mDeBERTa-v3-base。研究了多任务学习、基于翻译的数据增强及类别加权等技术,以应对严重标签不平衡问题。实验表明,独立任务建模配合类别加权策略更具有效性。

原文摘要 · Abstract (English)

This paper presents our system for SemEval-2026 Task 9: Detecting Multilingual, Multicultural and Multievent Online Polarization, which identifies polarized social media content in 22 languages through three subtasks: binary detection, target classification, and manifestation identification. We propose a heterogeneous ensemble of multilingual pretrained models, combining XLM-RoBERTa-large and mDeBERTa-v3-base. We investigate techniques such as multi-task learning, translation-based data augmentation, and class weighting to improve classification performance under severe label imbalance. Our findings indicate that independent task modeling combined with class weighting is more effective.

多语言极化检测集成学习跨文化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。