arXiv:2412.13666cs.CLcs.AI2024-12ACL被引 28

测试大模型生成个性化假新闻的能力及安全机制失效问题

Evaluation of LLM Vulnerabilities to Being Misused for Personalized Disinformation Generation

  • 评估主流大模型生成个性化假新闻的可行性
  • 发现个性化内容降低安全过滤触发率,反而易绕过检测
  • 揭示现有安全机制形同虚设,需紧急修复

近期大型语言模型(LLMs)生成高质量、人类难以分辨的文本能力引发滥用担忧。已有研究证实,这些模型可有效生成符合预设叙事的假新闻。其个性化内容生成能力也已被验证可用,但二者结合尚未系统评估。本研究填补空白,测试了开放与闭源最新大模型在英文环境下生成个性化假新闻的能力,并探究其是否能可靠评估个性化质量,以及个性化是否影响文本可检测性。结果表明,多数模型的安全过滤机制未正常工作,且个性化实际降低了安全过滤的激活频率,起到了‘越狱’效果。该现象必须由模型开发者和服务提供商立即应对。

原文摘要 · Abstract (English)

The capabilities of recent large language models (LLMs) to generate high-quality content indistinguishable by humans from human-written texts raises many concerns regarding their misuse. Previous research has shown that LLMs can be effectively misused for generating disinformation news articles following predefined narratives. Their capabilities to generate personalized (in various aspects) content have also been evaluated and mostly found usable. However, a combination of personalization and disinformation abilities of LLMs has not been comprehensively studied yet. Such a dangerous combination should trigger integrated safety filters of the LLMs, if there are some. This study fills this gap by evaluating vulnerabilities of recent open and closed LLMs, and their willingness to generate personalized disinformation news articles in English. We further explore whether the LLMs can reliably meta-evaluate the personalization quality and whether the personalization affects the generated-texts detectability. Our results demonstrate the need for stronger safety-filters and disclaimers, as those are not properly functioning in most of the evaluated LLMs. Additionally, our study revealed that the personalization actually reduces the safety-filter activations; thus effectively functioning as a jailbreak. Such behavior must be urgently addressed by LLM developers and service providers.

大模型安全假新闻生成个性化越狱攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。