arXiv:2503.24206cs.CL2025-03被引 2

用大模型生成假新闻数据,提升识别能力

Synthetic News Generation for Fake News Classification

  • 基于真实新闻改写关键事实,生成连贯假新闻
  • BERT等模型用少量合成数据就可提升检测准确率
  • 聚焦事实矛盾的特征对识别假新闻最有效

本研究探索利用大语言模型(LLMs)通过基于事实的篡改手段生成合成假新闻,并提出一套评估指标:连贯性、差异性和正确性。方法上,从真实新闻中提取关键事实,进行修改后重新生成内容以模拟假新闻,同时保持语义连贯。实验对比了传统机器学习模型与基于Transformer的BERT等模型在使用合成数据时的表现,结果表明,即使合成数据比例较小,BERT等模型也能有效提升假新闻检测效果。此外,专注于识别事实矛盾的特征在区分合成假新闻方面表现最佳。研究证实合成数据能增强假新闻检测系统,为未来研究提供重要参考,指出优化合成数据生成策略可进一步强化检测模型。

原文摘要 · Abstract (English)

This study explores the generation and evaluation of synthetic fake news through fact based manipulations using large language models (LLMs). We introduce a novel methodology that extracts key facts from real articles, modifies them, and regenerates content to simulate fake news while maintaining coherence. To assess the quality of the generated content, we propose a set of evaluation metrics coherence, dissimilarity, and correctness. The research also investigates the application of synthetic data in fake news classification, comparing traditional machine learning models with transformer based models such as BERT. Our experiments demonstrate that transformer models, especially BERT, effectively leverage synthetic data for fake news detection, showing improvements with smaller proportions of synthetic data. Additionally, we find that fact verification features, which focus on identifying factual inconsistencies, provide the most promising results in distinguishing synthetic fake news. The study highlights the potential of synthetic data to enhance fake news detection systems, offering valuable insights for future research and suggesting that targeted improvements in synthetic data generation can further strengthen detection models.

假新闻生成大模型应用文本检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。