测试大模型在政治操控等社会危害请求下的脆弱性,发现其防护失效严重。
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
- 构建覆盖7类社会政治领域、34国的585个有害请求数据集
- Mistral-7B在历史修正等场景下攻击成功率高达97%~98%
- 模型对21世纪及拉美、美英等地的敏感议题最易失守
大型语言模型(LLMs)正被广泛应用于可能产生直接社会政治后果的场景。然而,现有安全基准很少测试模型在政治操纵、宣传与虚假信息生成、监控与信息控制等领域的漏洞。我们提出SocialHarmBench,一个包含585个提示的数据集,覆盖7个社会政治类别和34个国家,旨在揭示模型在政治敏感情境下的关键失败点。评估显示,开放权重模型在有害请求面前高度脆弱,如Mistral-7B在历史修正主义、宣传和政治操纵等场景中攻击成功率达97%至98%。时间与地理分析表明,模型在面对21世纪或20世纪前的背景,以及拉美、美国、英国等地区相关问题时最为脆弱。这些发现表明当前安全防护无法泛化至高风险社会政治场景,暴露了系统性偏见,引发对大模型维护人权与民主价值可靠性的担忧。数据集已公开于https://huggingface.co/datasets/psyonp/SocialHarmBench。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed in contexts where their failures can have direct sociopolitical consequences. Yet, existing safety benchmarks rarely test vulnerabilities in domains such as political manipulation, propaganda and disinformation generation, or surveillance and information control. We introduce SocialHarmBench, a dataset of 585 prompts spanning 7 sociopolitical categories and 34 countries, designed to surface where LLMs most acutely fail in politically charged contexts. Our evaluations reveal several shortcomings: open-weight models exhibit high vulnerability to harmful compliance, with Mistral-7B reaching attack success rates as high as 97% to 98% in domains such as historical revisionism, propaganda, and political manipulation. Moreover, temporal and geographic analyses show that LLMs are most fragile when confronted with 21st-century or pre-20th-century contexts, and when responding to prompts tied to regions such as Latin America, the USA, and the UK. These findings demonstrate that current safeguards fail to generalize to high-stakes sociopolitical settings, exposing systematic biases and raising concerns about the reliability of LLMs in preserving human rights and democratic values. We share the SocialHarmBench benchmark at https://huggingface.co/datasets/psyonp/SocialHarmBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。