arXiv:2603.20925cs.AI2026-03被引 1

用利润驱动的对手测试智能体在经济互动中的漏洞,发现隐藏攻击策略。

Profit is the Red Team: Stress-Testing Agents in Strategic Economic Interactions

  • 用可学习对手最大化利润来替代人工构造攻击
  • 对手自动发现探查、锚定、欺骗承诺等攻击策略
  • 生成简洁提示规则,显著提升智能体抗压能力

随着智能体系统进入现实应用,其决策越来越依赖外部输入,如检索内容、工具输出和其他参与者提供的信息。当这些输入可能被对手战略性操纵时,安全风险已超越固定提示攻击,扩展至能引导智能体走向不利结果的自适应策略。我们提出利润驱动的红队测试协议,用一个仅通过标量结果反馈优化利润的学习对手替代手工攻击。该方法无需LLM作为裁判评分、攻击标签或攻击分类体系,适用于有可审计结果的结构化场景。我们在四个经典经济交互的简化环境中实现该协议,提供受控测试平台以评估自适应可利用性。控制实验表明,看似强大的智能体在利润优化对手压力下均表现出持续可被利用性,且学习对手自发发现探查、锚定和欺骗性承诺,无需显式指令。随后,我们将这些可利用事件提炼为简洁提示规则,使大多数先前观察到的失败失效,并显著提升目标性能。结果表明,利润驱动的红队数据可为具有可审计结果的结构化智能体设置提供实用的鲁棒性提升路径。

原文摘要 · Abstract (English)

As agentic systems move into real-world deployments, their decisions increasingly depend on external inputs such as retrieved content, tool outputs, and information provided by other actors. When these inputs can be strategically shaped by adversaries, the relevant security risk extends beyond a fixed library of prompt attacks to adaptive strategies that steer agents toward unfavorable outcomes. We propose profit-driven red teaming, a stress-testing protocol that replaces handcrafted attacks with a learned opponent trained to maximize its profit using only scalar outcome feedback. The protocol requires no LLM-as-judge scoring, attack labels, or attack taxonomy, and is designed for structured settings with auditable outcomes. We instantiate it in a lean arena of four canonical economic interactions, which provide a controlled testbed for adaptive exploitability. In controlled experiments, agents that appear strong against static baselines become consistently exploitable under profit-optimized pressure, and the learned opponent discovers probing, anchoring, and deceptive commitments without explicit instruction. We then distill exploit episodes into concise prompt rules for the agent, which make most previously observed failures ineffective and substantially improve target performance. These results suggest that profit-driven red-team data can provide a practical route to improving robustness in structured agent settings with auditable outcomes.

智能体安全红队测试经济博弈对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。