攻击者可微调图像,让模型误判非训练数据为训练数据,威胁隐私。
A Unified Perspective on Adversarial Membership Manipulation in Vision Models
- 通过微小扰动伪造成员身份,使非训练图像被误判为训练数据。
- 发现梯度范数坍缩轨迹是伪造与真实成员的关键区分特征。
- 提出基于梯度几何的检测方法,显著提升模型对恶意操纵的防御力。
成员推断攻击(MIAs)旨在判断特定数据是否属于模型的训练集,是评估视觉模型隐私泄露的有效工具。然而现有MIAs隐含假设查询输入为诚实数据,其对抗鲁棒性尚未被探索。本文首次揭示视觉模型中的对抗性成员操纵现象:微小且不可察觉的扰动可稳定地将非成员图像推入先进MIAs的‘成员’区域。我们提供了首个统一视角,分析该现象机制与影响。实验表明,对抗性成员伪造在多种架构与数据集上均有效。进一步发现一种独特的几何特征——梯度范数坍缩轨迹,能可靠区分伪造成员与真实成员,尽管二者语义几乎相同。基于此,我们提出基于梯度几何信号的检测策略,构建了鲁棒推理框架,显著提升抗操纵能力。大量实验证明,伪造普遍有效,而所提方法显著增强模型韧性。本工作建立了视觉模型中对抗性成员操纵的首个综合框架。
原文摘要 · Abstract (English)
Membership inference attacks (MIAs) aim to determine whether a specific data point was part of a model's training set, serving as effective tools for evaluating privacy leakage of vision models. However, existing MIAs implicitly assume honest query inputs, and their adversarial robustness remains unexplored. We show that MIAs for vision models expose a previously overlooked adversarial surface: adversarial membership manipulation, where imperceptible perturbations can reliably push non-member images into the "member" region of state-of-the-art MIAs. In this paper, we provide the first unified perspective on this phenomenon by analyzing its mechanism and implications. We begin by demonstrating that adversarial membership fabrication is consistently effective across diverse architectures and datasets. We then reveal a distinctive geometric signature - a characteristic gradient-norm collapse trajectory - that reliably separates fabricated from true members despite their nearly identical semantic representations. Building on this insight, we introduce a principled detection strategy grounded in gradient-geometry signals and develop a robust inference framework that substantially mitigates adversarial manipulation. Extensive experiments show that fabrication is broadly effective, while our detection and robust inference strategies significantly enhance resilience. This work establishes the first comprehensive framework for adversarial membership manipulation in vision models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。