arXiv:2605.02712cs.CLcs.AI2026-05ACL被引 1

用数据增强和自训练提升小样本下的阴谋论检测效果

mdok-style at SemEval-2026 Task 10: Finetuning LLMs for Conspiracy Detection

  • 基于数据增强与自训练,微调Qwen3-32B模型进行二分类
  • 在52个提交中排名前8(85百分位),表现强劲
  • 源自机器生成文本检测的方法首次有效应用于阴谋论识别

SemEval-2026 Task 10聚焦于阴谋论检测,目标是判断Reddit评论是否表达阴谋信念。我们提交的mdok-style系统采用数据增强与自训练策略,应对训练数据量较小的问题,对Qwen3-32B模型进行微调,完成二分类任务。该系统表现优异,在52个参赛方案中位列第8,处于85百分位。结果表明,最初用于机器生成文本检测的方法同样适用于阴谋论识别。

原文摘要 · Abstract (English)

SemEval-2026 Task 10 is focused on conspiracy detection. Specifically, the goal is to detect whether a Reddit comment expresses a conspiracy belief. Our submitted mdok-style system utilizes data augmentation and self-training (to cope with a rather small amount of training data) to finetune the Qwen3-32B model for a binary text-classification task. The submitted system is very competitive, ranking in the 85th percentile (8th out of 52 submissions). The results shown that our approach, which originated in machine-generated text detection, can be used for conspiracy detection as well.

阴谋论检测大模型微调自训练数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。