arXiv:2604.01014cs.CRcs.CV2026-04

用智能体自动探索策略,提升模型成员推断攻击效果

AutoMIA: Improved Baselines for Membership Inference Attack via Agentic Self-Exploration

  • 构建智能体框架,自主生成并优化攻击策略
  • 在多个大模型上表现优于现有基线,无需人工特征工程
  • 适合关注模型隐私泄露评估的研究者使用

成员推断攻击(MIA)是评估机器学习模型训练数据泄露的关键工具。然而,现有方法多依赖静态、手工设计的启发式规则,缺乏适应性,跨不同大模型迁移时性能往往不佳。本文提出 AutoMIA,一种基于智能体的框架,将成员推断重构为自动化自我探索与策略演化过程。给定高层次场景描述后,AutoMIA 通过生成可执行的对数概率级策略,并利用闭环评估反馈逐步优化。通过解耦抽象策略推理与底层执行,该框架实现了对攻击搜索空间的系统性、模型无关遍历。大量实验表明,AutoMIA 在多个大模型上均达到或超越当前最优基线,且无需手动特征工程。

原文摘要 · Abstract (English)

Membership Inference Attacks (MIAs) serve as a fundamental auditing tool for evaluating training data leakage in machine learning models. However, existing methodologies predominantly rely on static, handcrafted heuristics that lack adaptability, often leading to suboptimal performance when transferred across different large models. In this work, we propose AutoMIA, an agentic framework that reformulates membership inference as an automated process of self-exploration and strategy evolution. Given high-level scenario specifications, AutoMIA self-explores the attack space by generating executable logits-level strategies and progressively refining them through closed-loop evaluation feedback. By decoupling abstract strategy reasoning from low-level execution, our framework enables a systematic, model-agnostic traversal of the attack search space. Extensive experiments demonstrate that AutoMIA consistently matches or outperforms state-of-the-art baselines while eliminating the need for manual feature engineering.

隐私攻击智能体模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。