arXiv:2409.05208cs.LGcs.AI2024-09被引 4

研究发现,基于影响函数的归因结果可被恶意操纵,威胁数据价值评估可靠性。

Influence-based Attributions can be Manipulated

  • 利用反向传播友好方法设计高效攻击,系统性篡改归因结果。
  • 在ResNet特征与标准公平数据集上成功实现操纵,证实风险真实存在。
  • 适用于关注模型可解释性安全的研究者与实际应用开发者。

影响函数是用于将模型预测归因于训练数据的标准化工具,广泛应用于数据估值和公平性分析。本文揭示了操纵基于影响函数的归因结果的现实动机,并探究其是否可在对抗环境下被系统性篡改。研究发现,对于在ResNet特征嵌入上训练的逻辑回归模型及标准表格公平性数据集,此类操纵确实可行,并提出了具有反向传播兼容性的高效攻击方法。本工作质疑了影响函数在对抗场景下的可靠性。代码已公开于:https://github.com/infinite-pursuits/influence-based-attributions-can-be-manipulated。

原文摘要 · Abstract (English)

Influence Functions are a standard tool for attributing predictions to training data in a principled manner and are widely used in applications such as data valuation and fairness. In this work, we present realistic incentives to manipulate influence-based attributions and investigate whether these attributions can be \textit{systematically} tampered by an adversary. We show that this is indeed possible for logistic regression models trained on ResNet feature embeddings and standard tabular fairness datasets and provide efficient attacks with backward-friendly implementations. Our work raises questions on the reliability of influence-based attributions in adversarial circumstances. Code is available at : \url{https://github.com/infinite-pursuits/influence-based-attributions-can-be-manipulated}

可解释性对抗攻击数据估值

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。