arXiv:2511.11842cs.LGcs.CR2025-11

透明度越高,越易受攻击,防御者隐藏模型状态反而更安全

On the Trade-Off Between Transparency and Security in Adversarial Machine Learning

  • 用博弈论分析攻击者与防御者选择模型时的策略互动
  • 实测9种攻击、181个模型,匹配对方决策可提升攻击成功率
  • 揭示透明性与安全性本质矛盾,适合关注AI安全的研究者

透明性与安全性是负责任AI的核心,但在对抗场景下可能冲突。本文通过可迁移对抗样本攻击研究双方策略:攻击者利用替代模型扰动输入以欺骗目标模型,而防御者可选择是否防御。基于对9种攻击、181个模型的大规模实证评估,发现当攻击者匹配防御者的决策(是否防御)时成功率更高,表明隐蔽性对防御者有利。通过构建纳什博弈与斯塔克尔伯格博弈模型,分析其预期结果,证实仅知晓防御模型是否被保护就足以削弱安全。该结果揭示了透明性与安全性之间的普遍权衡,说明在对抗环境中,过度透明可能带来风险。本研究还展示了博弈论在识别透明性与安全冲突中的应用价值。

原文摘要 · Abstract (English)

Transparency and security are both central to Responsible AI, but they may conflict in adversarial settings. We investigate the strategic effect of transparency for agents through the lens of transferable adversarial example attacks. In transferable adversarial example attacks, attackers maliciously perturb their inputs using surrogate models to fool a defender's target model. These models can be defended or undefended, with both players having to decide which to use. Using a large-scale empirical evaluation of nine attacks across 181 models, we find that attackers are more successful when they match the defender's decision; hence, obscurity could be beneficial to the defender. With game theory, we analyze this trade-off between transparency and security by modeling this problem as both a Nash game and a Stackelberg game, and comparing the expected outcomes. Our analysis confirms that only knowing whether a defender's model is defended or not can sometimes be enough to damage its security. This result serves as an indicator of the general trade-off between transparency and security, suggesting that transparency in AI systems can be at odds with security. Beyond adversarial machine learning, our work illustrates how game-theoretic reasoning can uncover conflicts between transparency and security.

对抗机器学习博弈论模型安全透明性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。