构建大规模伪造说服攻击数据集,助力组织防范AI生成的舆论威胁。
Building Resilient Information Ecosystems: Large LLM-Generated Dataset of Persuasion Attacks
- 用GPT-4、Gemma 2和Llama 3.1生成13.4万条针对官方新闻的说服攻击。
- 发现不同模型偏好不同道德维度:GPT-4重关怀,Llama 3.1重忠诚。
- 数据集含两种传播形式,支持研究与防御虚假信息策略。
机构沟通对公共信任至关重要,但生成式AI模型可高速大规模制造具有说服力的内容,形成与政府及商业机构官方信息相竞争的叙事。这使机构处于被动应对状态,且常无法了解这些模型如何构建说服策略,从而削弱沟通效力。本文提出一个由GPT-4、Gemma 2和Llama 3.1生成的大型说服攻击数据集,涵盖134,136条攻击,覆盖SemEval 2023 Task 3中的23种说服技巧,针对10个机构的972份新闻稿。攻击内容以新闻稿陈述和社交媒体帖子两种形式呈现,包含长文本与短文本策略。分析显示,GPT-4攻击主要聚焦于‘关怀’(Care),同时涉及‘权威’(Authority)与‘忠诚’(Loyalty);Gemma 2强调‘关怀’与‘权威’;而Llama 3.1则侧重‘忠诚’与‘关怀’。该数据集可用于主动防御研究,帮助机构构建声誉防护机制,推动信息生态中更有效、更稳健的沟通发展。
原文摘要 · Abstract (English)
Organization's communication is essential for public trust, but the rise of generative AI models has introduced significant challenges by generating persuasive content that can form competing narratives with official messages from government and commercial organizations at speed and scale. This has left agencies in a reactive position, often unaware of how these models construct their persuasive strategies, making it more difficult to sustain communication effectiveness. In this paper, we introduce a large LLM-generated persuasion attack dataset, which includes 134,136 attacks generated by GPT-4, Gemma 2, and Llama 3.1 on agency news. These attacks span 23 persuasive techniques from SemEval 2023 Task 3, directed toward 972 press releases from ten agencies. The generated attacks come in two mediums, press release statements and social media posts, covering both long-form and short-form communication strategies. We analyzed the moral resonance of these persuasion attacks to understand their attack vectors. GPT-4's attacks mainly focus on Care, with Authority and Loyalty also playing a role. Gemma 2 emphasizes Care and Authority, while Llama 3.1 centers on Loyalty and Care. Analyzing LLM-generated persuasive attacks across models will enable proactive defense, allow to create the reputation armor for organizations, and propel the development of both effective and resilient communications in the information ecosystem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。