提出新攻击方法,让指代表达分割模型在多种描述下失效
Proxy-Embedding as an Adversarial Teacher: An Embedding-Guided Bidirectional Attack for Referring Expression Segmentation Models
- 用嵌入引导双向攻击,模拟多语言描述下的对抗样本
- 在多个主流模型和数据集上攻击成功率超基线15%以上
- 适合研究模型安全与隐私保护的开发者使用
指代表达分割(RES)可根据自然语言描述精准分割图像中的目标,具有高灵活性和广泛的应用前景。尽管性能出色,其对对抗样本的鲁棒性仍缺乏深入研究。现有攻击方法在常规分割模型上表现良好,但直接应用于多模态的RES模型时效果差,无法有效暴露其结构漏洞。在真实开放场景中,用户常对同一图像给出多种不同描述,因此需要能跨文本泛化的对抗样本。此外,从隐私保护角度出发,确保模型不会在未授权情况下分割敏感内容,是提升多模态视觉-语言系统安全性的关键。为此,我们提出嵌入引导双向攻击方法(PEAT),在多个主流RES架构和标准数据集上的大量实验表明,PEAT显著优于现有基线方法。
原文摘要 · Abstract (English)
Referring Expression Segmentation (RES) enables precise object segmentation in images based on natural language descriptions, offering high flexibility and broad applicability in real-world vision tasks. Despite its impressive performance, the robustness of RES models against adversarial examples remains largely unexplored. While prior adversarial attack methods have explored adversarial robustness on conventional segmentation models, they perform poorly when directly applied to RES models, failing to expose vulnerabilities in its multimodal structure. In practical open-world scenarios, users typically issue multiple, diverse referring expressions to interact with the same image, highlighting the need for adversarial examples that generalize across varied textual inputs. Furthermore, from the perspective of privacy protection, ensuring that RES models do not segment sensitive content without explicit authorization is a crucial aspect of enhancing the robustness and security of multimodal vision-language systems. To address these challenges, we present PEAT, an Embedding-Guided Bidirectional Attack for RES models. Extensive experiments across multiple RES architectures and standard benchmarks show that PEAT consistently outperforms competitive baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。