攻击者可隐秘操控模型公平性与解释结果,让敏感特征失效
The Unseen Hand: Manipulating Model Fairness and SHAP with Targeted Identity Re-Association Attacks

- 通过局部随机置换和排序扰动,无声无息改写模型输出
- 使公平性指标趋近理想值,且对保护特征的解释归因几乎为零
- 无需模型内部信息,适合隐蔽攻击,威胁现有审计机制
随着机器学习模型影响力扩大,算法公平性和可解释性成为问责关键。然而我们发现,这些审计机制本身易受隐蔽操纵,掩盖敏感特征的影响。尽管已有无需数据的攻击方法揭示此漏洞,但会留下可检测痕迹。本文提出目标身份重关联(TIRA)攻击,一种无需访问模型内部或特征表示、通过迭代概率方式操控输出的新方法。提出两种算法:概率微调换(PMiS)进行局部邻近交换,概率秩移微扰(PRSMP)引入随机小秩移。实验证明TIRA能有效将公平性指标推向理想值。关键突破在于成功混淆基于SHAP的解释,使保护特征的残余归因几乎为零,显著优于以往工作。
原文摘要 · Abstract (English)
As machine learning models grow more influential and opaque, algorithmic fairness and explainability are critical for ensuring accountability. However, we demonstrate that these auditing mechanisms are themselves vulnerable to subtle manipulation, camouflaging the influence of protected features. While prior work on data-agnostic attacks has exposed this vulnerability, they leave behind detectable artifacts that compromise their stealth. We introduce Targeted Identity Re-Association (TIRA) attacks, a novel family of attacks that iteratively and probabilistically manipulate a model's outputs without requiring access to the model's internals or feature representations. We formalize two algorithms: Probabilistic Micro-Shuffling (PMiS), which applies localized adjacent swaps, and Probabilistic Rank-Shift Micro-Perturbation (PRSMP), which introduces small, randomized rank shifts. We empirically demonstrate that TIRA attacks are highly effective at pushing fairness metrics towards ideal values. Crucially, TIRA attacks successfully confound SHAP-based explanations, leaving effectively zero residual attribution for protected features, a major improvement over prior work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。