用自研模型构建虚假信息传播环境,揭示大模型放大谣言的风险。
LLM Echo Chamber: personalized and automated disinformation
- 搭建模拟社交聊天室的虚假信息传播实验环境
- 经GPT4评估,生成内容具有高说服力与危害性
- 警示大模型在舆论引导中的伦理风险,适合安全研究者参考
大型语言模型(如GPT4、Llama2)在摘要、翻译和内容审核等任务中表现卓越,但其广泛应用引发担忧:模型可能大规模生成类人可信的误导性信息,显著影响公众舆论。本研究聚焦于大模型传播虚假信息为事实的风险,构建了名为LLM Echo Chamber的受控数字环境,模拟社交媒体聊天室中信息传播场景。该环境利用微软phi2模型,基于自定义数据集进行微调,生成有害内容。通过GPT4评估其说服力与危害性,结果显示生成内容极具误导性。研究揭示了大模型在信息操纵中的潜在威胁,强调需加强防范机制。
原文摘要 · Abstract (English)
Recent advancements have showcased the capabilities of Large Language Models like GPT4 and Llama2 in tasks such as summarization, translation, and content review. However, their widespread use raises concerns, particularly around the potential for LLMs to spread persuasive, humanlike misinformation at scale, which could significantly influence public opinion. This study examines these risks, focusing on LLMs ability to propagate misinformation as factual. To investigate this, we built the LLM Echo Chamber, a controlled digital environment simulating social media chatrooms, where misinformation often spreads. Echo chambers, where individuals only interact with like minded people, further entrench beliefs. By studying malicious bots spreading misinformation in this environment, we can better understand this phenomenon. We reviewed current LLMs, explored misinformation risks, and applied sota finetuning techniques. Using Microsoft phi2 model, finetuned with our custom dataset, we generated harmful content to create the Echo Chamber. This setup, evaluated by GPT4 for persuasiveness and harmfulness, sheds light on the ethical concerns surrounding LLMs and emphasizes the need for stronger safeguards against misinformation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。