arXiv:2606.26356cs.AIcs.IR2026-06中稿 · ICML

提示模块间存在隐性干扰,影响系统整体行为。

Instruction Bleed: Cross-Module Interference in Prompt-Composed Agentic Systems

  • 通过内容扰动实验发现模块间行为泄露
  • 内容扰动导致效应显著(Cohen's d=0.63)
  • 适合关注提示工程可靠性的研究者

使用提示组合的智能体系统常出现一种隐蔽故障:修改一个提示模块会悄然改变其他模块的行为,尽管它们无共享变量或执行依赖。我们将其定义为组合行为泄漏(CBL):共享上下文窗口的模块之间因架构非隔离而产生干扰。基于部署的职位评估智能体(Claude Sonnet 4.6,144次试验),我们提出可复用的三通道协议,分别扰动非焦点模块的体积、内容和形式。仅内容通道产生可检测的配对效应(Cohen's d = 0.63,Bootstrap 95% CI 排除零值);未发生推荐翻转——处于标准QA检测阈值以下,但会在数千次决策中累积。CBL与已知智能体失效机制(对抗注入、认知退化、多智能体故障传播、隐私泄露)正交。我们贡献了操作定义、可复用协议、可检验预测集及系统级表征,确立跨模块干扰测量是提示组合智能体评估的必要要求。

原文摘要 · Abstract (English)

Practitioners of prompt-composed agentic systems report a recurring failure mode: editing one prompt module silently shifts the behavior of others despite no shared variable or executable dependency. We formalize this as compositional behavioral leakage (CBL): interference between modules sharing a context window. CBL is enabled by architectural non-isolation: transformer self-attention provides no formal boundary between concatenated modules. We probe CBL on a deployed job-evaluation agent (Claude Sonnet 4.6, 144 trials) through a reusable three-channel protocol that perturbs non-focal modules along volume, content, and form. Only the content channel produces a detectable paired effect (Cohen's d = 0.63, bootstrap 95% CI excluding zero); no recommendation flipped -- a sub-threshold regime invisible to standard QA but compounding across the thousands of decisions a deployed agent makes. CBL is orthogonal to known agent-failure axes (adversarial injection, cognitive degradation, multi-agent fault propagation, privacy leakage). We contribute an operational definition, a reusable protocol, a falsifiable prediction set, and a system-class characterization, establishing cross-module interference measurement as a requirement for prompt-composed agent evaluation.

提示工程智能体系统行为泄漏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。