arXiv:2512.03001cs.AI2025-12

通过插入控制语句,防止大模型在长上下文中的越狱行为。

Invasive Context Engineering to Control Large Language Models

  • 在上下文中植入控制句子,实现无需训练的模型安全控制。
  • 实验证明该方法可有效降低长文本场景下的越狱概率。
  • 适用于需要高安全性且无法大量训练数据的场景。

当前研究通过偏好样本训练、提示工程和输入输出过滤来提升大语言模型对对抗攻击和异常行为的鲁棒性。尽管效果良好,大模型在长上下文场景中仍易被滥用,越狱概率随上下文长度增加而上升。亟需在长上下文情境下提供可靠的模型安全保证。本文提出一种侵入式上下文工程(Invasive Context Engineering),通过在模型输入上下文中插入控制语句,部分解决该问题。该方法可推广至思维链(Chain-of-Thought)过程,防止模型产生恶意策略。该技术不依赖模型训练,避免了长上下文训练中常见的数据短缺问题。

原文摘要 · Abstract (English)

Current research on operator control of Large Language Models improves model robustness against adversarial attacks and misbehavior by training on preference examples, prompting, and input/output filtering. Despite good results, LLMs remain susceptible to abuse, and jailbreak probability increases with context length. There is a need for robust LLM security guarantees in long-context situations. We propose control sentences inserted into the LLM context as invasive context engineering to partially solve the problem. We suggest this technique can be generalized to the Chain-of-Thought process to prevent scheming. Invasive Context Engineering does not rely on LLM training, avoiding data shortage pitfalls which arise in training models for long context situations.

大模型安全越狱防御上下文控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。