arXiv:2602.20708cs.AIcs.CR2026-02被引 4

用隐空间检测与修正,防住大模型代理的间接提示注入攻击。

ICON: Indirect Prompt Injection Defense for Agents based on Inference-Time Correction

  • 通过隐空间探针识别攻击导致的过度聚焦特征
  • 攻击成功率仅0.4%,任务可用性提升超50%
  • 适合需安全与效率平衡的智能代理系统

大型语言模型(LLM)代理易受间接提示注入(IPI)攻击,恶意检索内容中的指令会劫持代理执行流程。现有防御多依赖严格过滤或拒绝机制,存在过度拒绝问题,提前终止有效任务流程。本文提出ICON,一种基于推理时修正的探测-防御框架,可在保持任务连续性的同时中和攻击。核心洞察是IPI攻击会在隐空间留下明显的过度聚焦痕迹。我们引入隐空间痕迹探针,通过高强度得分检测攻击;随后由缓解校正器执行精准注意力引导,选择性调控对抗性查询-键依赖关系,同时增强任务相关元素,恢复LLM的正常执行轨迹。在多个骨干模型上的广泛评估表明,ICON实现0.4%的攻击成功率(ASR),媲美商用级检测器,同时任务实用性提升超过50%。此外,ICON展现出强泛化能力,可有效扩展至多模态代理,在安全与效率间取得更优平衡。

原文摘要 · Abstract (English)

Large Language Model (LLM) agents are susceptible to Indirect Prompt Injection (IPI) attacks, where malicious instructions in retrieved content hijack the agent's execution. Existing defenses typically rely on strict filtering or refusal mechanisms, which suffer from a critical limitation: over-refusal, prematurely terminating valid agentic workflows. We propose ICON, a probing-to-mitigation framework that neutralizes attacks while preserving task continuity. Our key insight is that IPI attacks leave distinct over-focusing signatures in the latent space. We introduce a Latent Space Trace Prober to detect attacks based on high intensity scores. Subsequently, a Mitigating Rectifier performs surgical attention steering that selectively manipulate adversarial query key dependencies while amplifying task relevant elements to restore the LLM's functional trajectory. Extensive evaluations on multiple backbones show that ICON achieves a competitive 0.4% ASR, matching commercial grade detectors, while yielding a over 50% task utility gain. Furthermore, ICON demonstrates robust Out of Distribution(OOD) generalization and extends effectively to multi-modal agents, establishing a superior balance between security and efficiency.

大模型安全提示注入代理系统隐空间分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。