arXiv:2608.16834cs.CLcs.AI2026-08

通过微小文本线索操控AI行为,揭示模型安全新风险

Model Hypnosis: Strong control of AI via additive subliminal effects

论文配图:Model Hypnosis: Strong control of AI via additive subliminal effects
图 1 · 摘自论文原文
  • 用看似无关的提示词组合实现对AI的强控制
  • 不同规模模型均受控,且提示可跨模型迁移
  • 适合关注AI安全与可解释性的研究者阅读

我们揭示了一种名为模型催眠的现象:在提示中系统性地组合个体弱但看似无关的线索,即可强烈操控AI模型行为。该现象存在于多种模型家族和规模中,包括前沿推理模型,且催眠提示具有跨模型迁移能力。由于控制依赖于不易察觉的文本细节(如改写、拼写错误),模型催眠为人工智能安全带来全新挑战,并成为理解模型行为的一大障碍。

原文摘要 · Abstract (English)

We demonstrate that AI models are broadly susceptible to a phenomenon we call model hypnosis, in which individually weak and seemingly irrelevant cues in the prompt can be systematically combined to strongly control model behavior. Model hypnosis occurs across model families and scales, including in frontier reasoning models, and hypnotic prompts can transfer between models. Because the model is controlled by inconspicuous textual choices, such as paraphrases and typos, model hypnosis presents new challenges and avenues for AI safety, and is a major hurdle for AI interpretability.

AI安全提示工程模型控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。