arXiv:2604.19844cs.CVcs.AI2026-04被引 1

视觉欺骗威胁智能体系统安全,新方法分离感知与决策提升抗干扰能力。

If you're waiting for a sign... that might not be it! Mitigating Trust Boundary Confusion from Visual Injections on Vision-Language Agentic Systems

论文配图:If you're waiting for a sign... that might not be it! Mitigating Trust Boundary Confusion from Visual Injections on Vision-Language Agentic Systems
图 1 · 摘自论文原文
  • 通过分离感知与决策模块,动态评估视觉输入可靠性。
  • 7个主流视觉语言智能体在攻击下误判率超60%,仅1个保持稳定响应。
  • 适用于需高可信度的自动驾驶、机器人等实时决策场景。

基于大视觉语言模型(LVLM)的具身视觉语言智能体系统(VLAS)使AI能感知并推理真实环境。然而,交通灯等环境信号可能被恶意篡改为误导性视觉注入,覆盖用户意图造成安全风险。这种合法信号与欺骗信号的混淆导致信任边界模糊。我们构建了双意图数据集与评估框架,发现当前LVLM智能体无法可靠平衡此矛盾,或忽视有效信号,或误信有害信号。我们在多个具身环境中对7个LVLM智能体进行结构化与噪声型注入测试。为此,提出多智能体防御框架,将感知与决策分离,动态评估视觉输入可信度。该方法显著降低误导行为,保持正确响应,并在对抗扰动下提供鲁棒性保障。评估框架与资源已公开于https://anonymous.4open.science/r/Visual-Prompt-Inject。

原文摘要 · Abstract (English)

Recent advances in embodied Vision-Language Agentic Systems (VLAS), powered by large vision-language models (LVLMs), enable AI systems to perceive and reason over real-world scenes. Within this context, environmental signals such as traffic lights are essential in-band signals that can and should influence agent behavior. However, similar signals could also be crafted to operate as misleading visual injections, overriding user intent and posing security risks. This duality creates a fundamental challenge: agents must respond to legitimate environmental cues while remaining robust to misleading ones. We refer to this tension as trust boundary confusion. To study this behavior, we design a dual-intent dataset and evaluation framework, through which we show that current LVLM-based agents fail to reliably balance this trade-off, either ignoring useful signals or following harmful ones. We systematically evaluate 7 LVLM agents across multiple embodied settings under both structure-based and noise-based visual injections. To address these vulnerabilities, we propose a multi-agent defense framework that separates perception from decision-making to dynamically assess the reliability of visual inputs. Our approach significantly reduces misleading behaviors while preserving correct responses and provides robustness guarantees under adversarial perturbations. The code of the evaluation framework and artifacts are made available at https://anonymous.4open.science/r/Visual-Prompt-Inject.

视觉语言智能体系统安全防御对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。