高探针AUC不等于能检测恶意注入,需结合多重诊断验证。
When AUC 0.998 Is Not Enough: A Candidate Evaluation Protocol for Hidden-State Probes of Indirect Prompt Injection in Multimodal Computer-Use Agents
- 提出多步诊断流程,检验隐藏状态探针是否真正识别恶意内容。
- 在Qwen2.5-VL-7B上测试,即使AUC达0.998也未必代表有效检测。
- 适合关注AI安全评估的科研人员和系统审计者使用。
隐藏状态探针——对冻结的视觉语言模型内部激活值进行线性分类——已成为在多模态计算机操作代理发出错误动作前,检测间接提示注入(IPI)的有力工具。本文以单个骨干模型案例(Qwen2.5-VL-7B在Mind2Web上,教师强制回放)为研究对象,指出:仅凭干净样本与攻击样本之间的高探针AUC(0.998),不能单独作为恶意内容检测的证据。通过两种后验诊断——文本侧注入的成对构造标量基线,以及叠加表面的同步干扰匹配视觉控制——发现该高AUC结果仍可能源于非语义线索,无法无条件解读为恶意内容感知。因此,我们构建了一套候选控制集及报告启发规则,明确说明高清洁-攻击AUC所能支持和不能支持的结论。标签为注入表面存在,而非攻击成功;跨模型与基准的泛化能力仍属推测。
原文摘要 · Abstract (English)
Hidden-state probing -- a linear classifier on a frozen vision-language model's internal activations -- has emerged as an attractive evaluation tool for flagging indirect prompt injection (IPI) in multimodal computer-use agents before the agent emits a corrupted action. We argue, on a single-backbone cautionary case study (Qwen2.5-VL-7B on Mind2Web, teacher-forced replay), that a high probing AUC on a clean-vs-attack split is not, on its own, evidence of malicious-content detection. Two post-hoc diagnostics -- a paired-construction scalar baseline on text-side injections, and same-step nuisance-matched visual controls on the overlay surface -- do not license an unqualified malicious-content interpretation of the headline while leaving room for partly-semantic readings. We package the diagnostics as a candidate control set with reporting heuristics for what a high clean-vs-attack AUC does and does not license. Labels are injection-surface-present, not attack success; generalisation beyond this backbone and benchmark is a conjecture.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。