arXiv:2505.18916cs.CL2025-05AAAI被引 1

构建首个覆盖9语言的谣言立场分类数据集,助力多语言假消息分析。

SCRum-9: Multilingual Stance Classification over Rumours on Social Media

  • 构建涵盖9种语言的谣言立场标注数据集,链接2100个事实核查条目。
  • 验证合成数据可提升小模型性能,超越大模型零样本推理效果。
  • 发现模型预测常匹配人类第二选择,反映对模糊案例的共识倾向。

我们提出SCRum-9,目前最大规模的多语言谣言立场分类数据集,覆盖X平台上的7,516条推文,涉及9种语言。该数据集突破现有局限,包含2,100个经核实的声明,引入多位标注者提供的置信度标注以捕捉标注间与标注内差异。每语言至少由两名母语者标注,总耗时超405小时,报酬达8,150美元。本文进一步用五款大语言模型(LLMs)和两款多语言掩码语言模型(MLMs)在上下文学习(ICL)与微调设置下进行基准测试。创新性地探索多语言合成数据应用,表明即使性能较弱的LLM也可生成高质量合成数据,用于微调小型MLM,使其表现优于大模型的零样本ICL。最后,分析模型预测与人类不确定性的关系,发现模型更倾向于匹配标注者的第二选择,而非完全偏离人类判断。SCRum-9已公开发布,有望推动社交媒体中误导性叙事的多语言研究。

原文摘要 · Abstract (English)

We introduce SCRum-9, the largest multilingual Stance Classification dataset for Rumour analysis in 9 languages, containing 7,516 tweets from X. SCRum-9 goes beyond existing stance classification datasets by covering more languages, linking examples to more fact-checked claims (2.1k), and including confidence-related annotations from multiple annotators to account for intra- and inter-annotator variability. Annotations were made by at least two native speakers per language, totalling more than 405 hours of annotation and 8,150 dollars in compensation. Further, SCRum-9 is used to benchmark five large language models (LLMs) and two multilingual masked language models (MLMs) in In-Context Learning (ICL) and fine-tuning setups. This paper also innovates by exploring the use of multilingual synthetic data for rumour stance classification, showing that even LLMs with weak ICL performance can produce valuable synthetic data for fine-tuning small MLMs, enabling them to achieve higher performance than zero-shot ICL in LLMs. Finally, we examine the relationship between model predictions and human uncertainty on ambiguous cases finding that model predictions often match the second-choice labels assigned by annotators, rather than diverging entirely from human judgments. SCRum-9 is publicly released to the research community with potential to foster further research on multilingual analysis of misleading narratives on social media.

谣言检测多语言立场分类数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。