研究大模型如何生成操纵性言论及有效抑制方法
When Agents Persuade: Rhetoric Generation and Mitigation in LLMs
- 让大模型模拟宣传任务,分析其使用多种修辞手段
- 微调可显著降低生成操纵内容倾向,ORPO效果最佳
- 适合关注AI伦理与安全的研究者和开发者
尽管基于大语言模型的智能体在开放环境中具有广泛应用价值,但可能被用于生成操纵性内容。本研究通过设定宣传目标,测试大模型输出,并利用两个领域专用模型(一个用于区分宣传与非宣传文本,另一个用于检测修辞技巧如情绪化语言、恐惧诉求、旗帜挥舞、人身攻击)进行分析。结果表明,大模型在提示下会表现出明显的宣传行为,采用多种修辞策略。我们进一步探索了监督微调(SFT)、直接偏好优化(DPO)和奇偶比偏好优化(ORPO)等缓解方法,发现微调能显著降低其生成此类内容的倾向,其中ORPO表现最优。
原文摘要 · Abstract (English)
Despite their wide-ranging benefits, LLM-based agents deployed in open environments can be exploited to produce manipulative material. In this study, we task LLMs with propaganda objectives and analyze their outputs using two domain-specific models: one that classifies text as propaganda or non-propaganda, and another that detects rhetorical techniques of propaganda (e.g., loaded language, appeals to fear, flag-waving, name-calling). Our findings show that, when prompted, LLMs exhibit propagandistic behaviors and use a variety of rhetorical techniques in doing so. We also explore mitigation via Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and ORPO (Odds Ratio Preference Optimization). We find that fine-tuning significantly reduces their tendency to generate such content, with ORPO proving most effective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。