arXiv:2503.02328cs.CLcs.CY2025-03被引 7

用大模型生成假信息数据,对新冠谣言立场识别效果有限。

Limited Effectiveness of LLM-based Data Augmentation for COVID-19 Misinformation Stance Detection

  • 用大模型可控生成谣言数据,用于扩充训练集。
  • 相比传统方法,性能提升小且不稳定,因模型自带防护机制。
  • 适合研究谣言检测与生成的学者参考,代码数据已开源。

新兴疫情相关的虚假信息对社会构成严重威胁,亟需有效应对措施。立场检测(SD)可识别社交媒体帖子是否支持或反对误导性声明,是关键手段之一。本文在包含声明与对应推文的新冠谣言立场检测数据集上微调分类器,测试了基于大语言模型(LLMs)的可控虚假信息生成(CMG)作为数据增强的方法。尽管CMG有扩展训练数据的潜力,但实验表明其性能提升常不明显且不一致,主要归因于大模型内部的安全机制。研究团队已公开代码与数据集,以推动虚假信息检测与生成的进一步研究。

原文摘要 · Abstract (English)

Misinformation surrounding emerging outbreaks poses a serious societal threat, making robust countermeasures essential. One promising approach is stance detection (SD), which identifies whether social media posts support or oppose misleading claims. In this work, we finetune classifiers on COVID-19 misinformation SD datasets consisting of claims and corresponding tweets. Specifically, we test controllable misinformation generation (CMG) using large language models (LLMs) as a method for data augmentation. While CMG demonstrates the potential for expanding training datasets, our experiments reveal that performance gains over traditional augmentation methods are often minimal and inconsistent, primarily due to built-in safeguards within LLMs. We release our code and datasets to facilitate further research on misinformation detection and generation.

谣言检测大模型数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。