arXiv:2603.19375cs.CRcs.LG2026-03

用大模型自动发现数据泄露攻击新方法,效果提升显著。

Automated Membership Inference Attacks: Discovering MIA Signal Computations using LLM Agents

  • 用大模型代理自动探索攻击策略空间,无需人工设计
  • 在多个模型上实现最高0.18的AUC提升,优于现有方法
  • 适合安全研究者和隐私防护开发者参考

成员推理攻击(MIA)能够判断特定数据是否属于模型训练集,是评估机器学习系统信息泄露风险的重要框架。传统攻击设计依赖大量人工探索,效率低。本文提出AutoMIA框架,利用大语言模型(LLM)代理自动搜索并生成新型MIA信号计算方式,系统性地探索攻击策略空间。实验表明,AutoMIA可针对用户指定的目标模型与数据集,发现定制化的新攻击策略,相比现有方法,绝对AUC提升最高达0.18。该工作首次证明了LLM代理可作为高效且可扩展的MIA设计范式,达到当前最优性能,为未来隐私安全研究开辟新路径。

原文摘要 · Abstract (English)

Membership inference attacks (MIAs), which enable adversaries to determine whether specific data points were part of a model's training dataset, have emerged as an important framework to understand, assess, and quantify the potential information leakage associated with machine learning systems. Designing effective MIAs is a challenging task that usually requires extensive manual exploration of model behaviors to identify potential vulnerabilities. In this paper, we introduce AutoMIA -- a novel framework that leverages large language model (LLM) agents to automate the design and implementation of new MIA signal computations. By utilizing LLM agents, we can systematically explore a vast space of potential attack strategies, enabling the discovery of novel strategies. Our experiments demonstrate AutoMIA can successfully discover new MIAs that are specifically tailored to user-configured target model and dataset, resulting in improvements of up to 0.18 in absolute AUC over existing MIAs. This work provides the first demonstration that LLM agents can serve as an effective and scalable paradigm for designing and implementing MIAs with SOTA performance, opening up new avenues for future exploration.

成员推理大模型应用隐私安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。