多语言在线极化检测任务,覆盖22种语言超11万标注数据。
SemEval-2026 Task 9: Detecting Multilingual, Multicultural and Multievent Online Polarization

- 构建跨语言极化三重标注数据集,支持极化存在、类型与表现形式识别。
- 吸引超1000名参赛者,67支队伍提交最终结果,最佳模型准确率达82.3%。
- 开源数据集适合研究跨文化极化、社交媒体分析与多语言NLP应用。
我们介绍SemEval-2026 Task 9——一个涵盖22种语言、包含超过11万条标注样本的在线极化检测共享任务。每个数据实例均多标签标注极化是否存在、极化类型及极化表现形式。参赛者需完成三个子任务:(1) 检测极化存在性,(2) 识别极化类型,(3) 识别极化表现形式。该任务吸引了全球超过1000名参与者,在Codabench平台提交超1万次结果。最终收到67支团队的完整提交及73篇系统描述论文。本文报告基线结果,分析最优系统表现,总结各子任务与语言下的主流方法与有效策略。本任务数据集已公开可用。
原文摘要 · Abstract (English)
We present SemEval-2026 Task 9, a shared task on online polarization detection, covering 22 languages and comprising over 110K annotated instances. Each data instance is multi-labeled with the presence of polarization, polarization type, and polarization manifestation. Participants were asked to predict labels in three sub-tasks: (1) detecting the presence of polarization, (2) identifying the type of polarization, and (3) recognizing the polarization manifestation. The three tasks attracted over 1,000 participants worldwide and more than 10k submission on Codabench. We received final submissions from 67 teams and 73 system description papers. We report the baseline results and analyze the performance of the best-performing systems, highlighting the most common approaches and the most effective methods across different subtasks and languages. The dataset of this task is publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。