用AI检测说服攻击并评估防护效果,发现大模型差异显著。
Proactive Defense: Compound AI for Detecting Persuasion Attacks and Measuring Inoculation Effectiveness
- 设计四类AI代理协同识别、生成、防御和评估说服攻击。
- GPT-4对复杂话术检测准确率最高,开源模型在微妙修辞上表现弱。
- 提示工程影响检测效果,不同模型需调参适配,适合安全与认知防护研究者。
本文提出BRIES,一种新型复合AI架构,用于检测信息环境中的说服攻击并评估防护效果。系统包含四类专用代理:生成特定说服策略的攻击生成器(Twister)、可配置参数的攻击类型识别器(Detector)、通过内容免疫机制生成抗性内容的防御者(Defender),以及利用因果推断评估免疫效果的评估者(Assessor)。在合成说服数据集上基于SemEval 2023 Task 3分类体系进行实验,结果显示语言模型检测性能存在显著差异:GPT-4在复杂说服技术上表现优异,而Llama3和Mistral在识别细微修辞策略时表现较弱,表明不同架构对说服语言模式的编码与处理方式本质不同。提示工程显著影响检测效能,温度设置与置信度评分导致模型特异性变化——Gemma和GPT-4在低温下表现更优,而Llama3和Mistral在高温下能力提升。因果分析揭示了说服攻击的社会情感认知特征,不同攻击类型针对特定认知维度。本研究推动生成式AI安全与认知安全发展,量化了大模型对说服攻击的脆弱性,并提供结构化干预框架以增强人类认知韧性。
原文摘要 · Abstract (English)
This paper introduces BRIES, a novel compound AI architecture designed to detect and measure the effectiveness of persuasion attacks across information environments. We present a system with specialized agents: a Twister that generates adversarial content employing targeted persuasion tactics, a Detector that identifies attack types with configurable parameters, a Defender that creates resilient content through content inoculation, and an Assessor that employs causal inference to evaluate inoculation effectiveness. Experimenting with the SemEval 2023 Task 3 taxonomy across the synthetic persuasion dataset, we demonstrate significant variations in detection performance across language agents. Our comparative analysis reveals significant performance disparities with GPT-4 achieving superior detection accuracy on complex persuasion techniques, while open-source models like Llama3 and Mistral demonstrated notable weaknesses in identifying subtle rhetorical, suggesting that different architectures encode and process persuasive language patterns in fundamentally different ways. We show that prompt engineering dramatically affects detection efficacy, with temperature settings and confidence scoring producing model-specific variations; Gemma and GPT-4 perform optimally at lower temperatures while Llama3 and Mistral show improved capabilities at higher temperatures. Our causal analysis provides novel insights into socio-emotional-cognitive signatures of persuasion attacks, revealing that different attack types target specific cognitive dimensions. This research advances generative AI safety and cognitive security by quantifying LLM-specific vulnerabilities to persuasion attacks and delivers a framework for enhancing human cognitive resilience through structured interventions before exposure to harmful content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。