用无痕后门检测数据泄露,隐蔽性强且成功率高。
Hide in Plain Sight: Clean-Label Backdoor for Auditing Membership Inference
- 用影子模型生成自然标签的隐蔽触发器。
- 黑盒攻击下在多个数据集上成功率达90%以上。
- 适合隐私审计与合规检查场景使用。
成员推理攻击(MIAs)是评估隐私风险、确保符合《通用数据保护条例》(GDPR)等法规的关键工具。然而,其在审计数据未经授权使用方面的潜力尚未充分挖掘。为此,我们提出一种基于干净标签后门的新方法,专为鲁棒且隐蔽的数据审计设计。与依赖可检测污染样本和标签变更的传统方法不同,该方法保留原始标签,即使在低污染率下也具备更强隐蔽性。通过影子模型生成最优触发器,模拟目标模型行为,使触发样本在特征空间中与源类别距离最小,同时保持原始标签不变。结果是一种强大且难以察觉的审计机制,克服了现有方法存在的标签不一致和视觉伪影等问题。该方法仅需黑盒访问,可在多种数据集和模型架构上实现高攻击成功率。此外,还解决了触发器隐蔽性和污染持久性挑战,成为实用高效的数据审计方案。全面实验验证了该方法的有效性和泛化能力,在隐蔽性和攻击成功率指标上均优于多个基线方法。
原文摘要 · Abstract (English)
Membership inference attacks (MIAs) are critical tools for assessing privacy risks and ensuring compliance with regulations like the General Data Protection Regulation (GDPR). However, their potential for auditing unauthorized use of data remains under explored. To bridge this gap, we propose a novel clean-label backdoor-based approach for MIAs, designed specifically for robust and stealthy data auditing. Unlike conventional methods that rely on detectable poisoned samples with altered labels, our approach retains natural labels, enhancing stealthiness even at low poisoning rates. Our approach employs an optimal trigger generated by a shadow model that mimics the target model's behavior. This design minimizes the feature-space distance between triggered samples and the source class while preserving the original data labels. The result is a powerful and undetectable auditing mechanism that overcomes limitations of existing approaches, such as label inconsistencies and visual artifacts in poisoned samples. The proposed method enables robust data auditing through black-box access, achieving high attack success rates across diverse datasets and model architectures. Additionally, it addresses challenges related to trigger stealthiness and poisoning durability, establishing itself as a practical and effective solution for data auditing. Comprehensive experiments validate the efficacy and generalizability of our approach, outperforming several baseline methods in both stealth and attack success metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。