小模型能生成以假乱真的假新闻标题,但质量检测效果差。
Do small language models generate realistic variable-quality fake news headlines?
- 用提示工程让14个小型模型生成真假新闻标题
- 检测模型准确率仅35.2%至63.5%,误判普遍
- 适合关注虚假信息风险与检测盲区的研究者
本研究评估了14个参数量在1.7B至14B之间的小型语言模型(包括LLaMA、Gemma、Phi、SmolLM、Mistral和Granite系列)在明确指令下生成低质量与高质量假新闻标题的能力,并检验其与真实新闻标题的相似性。通过控制提示工程,共生成24,000条标题,涵盖低质与高质误导类别。使用现有的基于DistilBERT和集成分类器的新闻标题质量检测模型对生成内容进行评估,结果显示:多数小模型表现出高合规性且伦理抵抗较弱,但存在个别例外;在质量检测任务中,误分类现象普遍,准确率仅为35.2%至63.5%。这表明所测试的小型语言模型总体上具备生成虚假标题的能力,尽管存在轻微伦理限制差异,但生成内容与网络上主流人工撰写的文本在质量特征上并不相近,因检测准确率偏低。
原文摘要 · Abstract (English)
Small language models (SLMs) have the capability for text generation and may potentially be used to generate falsified texts online. This study evaluates 14 SLMs (1.7B-14B parameters) including LLaMA, Gemma, Phi, SmolLM, Mistral, and Granite families in generating perceived low and high quality fake news headlines when explicitly prompted, and whether they appear to be similar to real-world news headlines. Using controlled prompt engineering, 24,000 headlines were generated across low-quality and high-quality deceptive categories. Existing machine learning and deep learning-based news headline quality detectors were then applied against these SLM-generated fake news headlines. SLMs demonstrated high compliance rates with minimal ethical resistance, though there were some occasional exceptions. Headline quality detection using established DistilBERT and bagging classifier models showed that quality misclassification was common, with detection accuracies only ranging from 35.2% to 63.5%. These findings suggest the following: tested SLMs generally are compliant in generating falsified headlines, although there are slight variations in ethical restraints, and the generated headlines did not closely resemble existing primarily human-written content on the web, given the low quality classification accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。