研究大模型如何被攻破生成假医疗信息,及如何用大模型识别这类信息。
An Audit and Analysis of LLM-Assisted Health Misinformation Jailbreaks Against LLMs
- 测试109种攻击方式,评估其让大模型生成有害医疗信息的效果。
- 发现大模型生成的假信息与社交媒体上的相似,难以区分。
- 证明大模型能有效检测来自其他大模型和人类的虚假信息。
大型语言模型(LLMs)既能无意中生成有害误导信息,也可能在‘越狱’攻击下被诱导输出恶意内容。本文研究了由大模型发起的越狱攻击在引发其他模型产生有害医疗误导信息方面的有效性与特征,并对比了这些攻击提示与真实社交平台上的健康相关查询。同时,分析了越狱后生成的误导信息与Reddit上常见健康类谣言的相似性,并评估了传统机器学习方法对这类信息的检测能力。研究共考察了针对三个目标模型的109种不同攻击方式,结果表明大模型在识别来自其他大模型及人类的虚假信息方面具有潜力,支持通过精心设计使大模型成为更健康信息生态的重要组成部分。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are a double-edged sword capable of generating harmful misinformation -- inadvertently, or when prompted by "jailbreak" attacks that attempt to produce malicious outputs. LLMs could, with additional research, be used to detect and prevent the spread of misinformation. In this paper, we investigate the efficacy and characteristics of LLM-produced jailbreak attacks that cause other models to produce harmful medical misinformation. We also study how misinformation generated by jailbroken LLMs compares to typical misinformation found on social media, and how effectively it can be detected using standard machine learning approaches. Specifically, we closely examine 109 distinct attacks against three target LLMs and compare the attack prompts to in-the-wild health-related LLM queries. We also examine the resulting jailbreak responses, comparing the generated misinformation to health-related misinformation on Reddit. Our findings add more evidence that LLMs can be effectively used to detect misinformation from both other LLMs and from people, and support a body of work suggesting that with careful design, LLMs can contribute to a healthier overall information ecosystem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。