首个跨语言跨地区假新闻生成攻击评估基准,揭示大模型安全防护漏洞。
JailNewsBench: Multi-Lingual and Regional Benchmark for Fake News Generation under Jailbreak Attacks
- 构建覆盖34地区22语言的多语言假新闻攻击测试集
- 最高攻击成功率86.3%,英语相关话题防御最弱
- 揭示现有安全数据集对假新闻覆盖不足,适合安全研究者使用
假新闻损害政治、经济、健康与国际关系中的社会信任与决策,极端情况下威胁生命与公共安全。由于假新闻反映特定区域的政治、社会与文化背景,并以语言形式呈现,评估大语言模型(LLMs)风险需具备多语言与区域性视角。恶意用户可通过越狱攻击绕过安全防护,诱导模型生成假新闻。然而,目前尚无系统性基准用于评估跨语言、跨区域的攻击鲁棒性。本文提出JailNewsBench,首个评估大模型在越狱攻击下生成假新闻鲁棒性的基准。该基准涵盖34个地区、22种语言,包含5种越狱攻击方式和8项评估子指标,共约30万条样本。对9个LLM的评估显示,最高攻击成功率(ASR)达86.3%,最高有害性得分3.5/5。值得注意的是,针对英语及美国相关话题,主流多语言模型的防御能力显著低于其他区域,凸显语言与地域间安全防护的严重不平衡。此外,分析表明现有安全数据集对假新闻的覆盖有限,且防御水平低于毒性与社会偏见等主要类别。数据集与代码已开源:https://github.com/kanekomasahiro/jail_news_bench。
原文摘要 · Abstract (English)
Fake news undermines societal trust and decision-making across politics, economics, health, and international relations, and in extreme cases threatens human lives and societal safety. Because fake news reflects region-specific political, social, and cultural contexts and is expressed in language, evaluating the risks of large language models (LLMs) requires a multi-lingual and regional perspective. Malicious users can bypass safeguards through jailbreak attacks, inducing LLMs to generate fake news. However, no benchmark currently exists to systematically assess attack resilience across languages and regions. Here, we propose JailNewsBench, the first benchmark for evaluating LLM robustness against jailbreak-induced fake news generation. JailNewsBench spans 34 regions and 22 languages, covering 8 evaluation sub-metrics through LLM-as-a-Judge and 5 jailbreak attacks, with approximately 300k instances. Our evaluation of 9 LLMs reveals that the maximum attack success rate (ASR) reached 86.3% and the maximum harmfulness score was 3.5 out of 5. Notably, for English and U.S.-related topics, the defensive performance of typical multi-lingual LLMs was significantly lower than for other regions, highlighting substantial imbalances in safety across languages and regions. In addition, our analysis shows that coverage of fake news in existing safety datasets is limited and less well defended than major categories such as toxicity and social bias. Our dataset and code are available at https://github.com/kanekomasahiro/jail_news_bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。