测试开源大模型在社交媒体中表达政治观点的能力及越狱效果
How Far Will They Go? Red-Teaming Online Influence with Large Language Models

- 通过实测30多个模型,评估其政治言论范围和越狱效果
- 模型越小越保守,左倾倾向明显,不同国家模型差异大
- 发现越狱技巧对不同模型效果差异显著,适合安全与舆情研究者
随着基于大语言模型(LLM)的智能体参与在线讨论,评估其支持政治影响力行动的能力对信息真实性至关重要。本文聚焦本地部署的开源大模型,因其更符合注重隐私的恶意行为者在社交媒体环境中的操作限制。提出一种实证红队测试框架,用于衡量大模型的奥弗顿窗口(OW),即模型在争议话题上可靠表达的政治观点范围,并量化自然语言越狱手段如何扩展该范围。评估了来自10个模型家族、5个不同国家的30多个模型。结果发现:政治表达存在系统性不对称——开源模型更倾向于生成左倾内容;奥弗顿窗口随模型规模增大而收缩;区域差异显著,尽管开源生态覆盖不均。越狱有效性在不同模型族间差异巨大,提示需优化越狱技术组合流程。整体成果为审计开源大模型的政治可操纵性提供了实用框架,并助力未来设计更强对抗性防御措施。
原文摘要 · Abstract (English)
As large language model (LLM)-based agents increasingly participate in online discourse, red-teaming their capacity to support political influence campaigns is critical for information integrity. In pursuit of this goal, we focus on locally deployed open-source LLMs, as opposed to frontier API-only models, given their superior alignment with the operational constraints of privacy-conscious malicious actors deployed in social media environments. We introduce an empirical red-teaming framework for measuring LLM Overton Windows (OWs), defined as the range of political opinions a model can reliably express on controversial topics, and for quantifying how simple natural-language jailbreaks expand that range. We evaluate more than 30 LLMs spanning 10 model families and five countries of origin. We find systematic asymmetries in political expressivity: open-source LLMs are typically more willing to generate left-leaning social media content, OWs tend to contract inversely to model size, and regional differences are substantial despite uneven representation in the open-source ecosystem. Jailbreak potency also varies sharply across model families, motivating a workflow for identifying effective combinations of jailbreak techniques. Taken together, our results establish a practical framework for auditing the political steerability of open-source LLMs and for helping future researchers design stronger countermeasures against LLM-enabled influence campaigns.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。