arXiv:2511.19536cs.CRcs.AI2025-11被引 2

用大模型自动发起机器学习服务推理攻击,无需专家经验。

AttackPilot: Autonomous Inference Attacks Against ML Services With LLM-Based Agents

  • 基于大模型构建自主攻击代理,可独立完成攻击任务。
  • 在20个目标服务上实现100%任务完成率,单次平均耗资仅0.627美元。
  • 适合非专家的模型服务商、审计方或监管者进行风险评估。

推理攻击已被广泛研究,能系统评估机器学习服务的风险;然而,其实施及最优攻击参数配置对非专家而言仍具挑战。大型语言模型的兴起为开发自主代理作为推理攻击专家提供了潜在但尚未充分探索的机会。本文提出AttackPilot,一种可独立执行推理攻击的自主代理,无需人工干预。我们在20个目标服务上对其进行了评估,结果表明,该代理使用GPT-4o时实现了100.0%的任务完成率,接近专家级攻击表现,单次运行平均令牌成本仅为0.627美元。该代理亦可由其他主流LLM驱动,并能在服务约束下自适应优化策略。我们还进行了追踪分析,证明多代理框架与任务特异性动作空间等设计有效缓解了错误计划、指令偏离、上下文丢失和幻觉等问题。我们预期此类代理可赋能非专家的模型服务提供方、审计机构或监管者,在无需深厚领域知识的前提下系统评估机器学习服务风险。

原文摘要 · Abstract (English)

Inference attacks have been widely studied and offer a systematic risk assessment of ML services; however, their implementation and the attack parameters for optimal estimation are challenging for non-experts. The emergence of advanced large language models presents a promising yet largely unexplored opportunity to develop autonomous agents as inference attack experts, helping address this challenge. In this paper, we propose AttackPilot, an autonomous agent capable of independently conducting inference attacks without human intervention. We evaluate it on 20 target services. The evaluation shows that our agent, using GPT-4o, achieves a 100.0% task completion rate and near-expert attack performance, with an average token cost of only $0.627 per run. The agent can also be powered by many other representative LLMs and can adaptively optimize its strategy under service constraints. We further perform trace analysis, demonstrating that design choices, such as a multi-agent framework and task-specific action spaces, effectively mitigate errors such as bad plans, inability to follow instructions, task context loss, and hallucinations. We anticipate that such agents could empower non-expert ML service providers, auditors, or regulators to systematically assess the risks of ML services without requiring deep domain expertise.

推理攻击大模型应用自动化安全模型审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。