arXiv:2504.02865cs.CLcs.LG2025-04被引 1

用语言巧设计,让大模型说出假话,暴露其事实漏洞。

The Illusionist's Prompt: Exposing the Factual Vulnerabilities of Large Language Models with Linguistic Nuances

  • 通过精心设计的语义微妙提示,诱导大模型产生事实错误。
  • 攻击在五种增强事实性的策略下仍有效,包括GPT-4o和Gemini-2.0。
  • 攻击保持用户意图不变,隐蔽性强,适合安全测试与防御研究。

随着大语言模型(LLMs)持续进步,非专业用户越来越依赖其作为实时信息来源。为保障信息真实性,现有研究多聚焦于正式查询中的幻觉问题,但忽视了恶意构造查询的威胁。本文提出「魔术师提示」(The Illusionist's Prompt),一种融合语言细微差异的对抗性攻击,挑战五类增强事实性的策略。该攻击可自动生成高度可迁移的误导性提示,引发模型内部事实错误,同时保留原始用户意图与语义。大量实验表明,该方法能有效攻破黑盒模型,包括GPT-4o和Gemini-2.0等商业API,即使面对多种防御机制依然奏效。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) continue to advance, they are increasingly relied upon as real-time sources of information by non-expert users. To ensure the factuality of the information they provide, much research has focused on mitigating hallucinations in LLM responses, but only in the context of formal user queries, rather than maliciously crafted ones. In this study, we introduce The Illusionist's Prompt, a novel hallucination attack that incorporates linguistic nuances into adversarial queries, challenging the factual accuracy of LLMs against five types of fact-enhancing strategies. Our attack automatically generates highly transferrable illusory prompts to induce internal factual errors, all while preserving user intent and semantics. Extensive experiments confirm the effectiveness of our attack in compromising black-box LLMs, including commercial APIs like GPT-4o and Gemini-2.0, even with various defensive mechanisms.

大模型安全幻觉攻击语言陷阱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。