针对遵守规则的AI客服,设计新对抗测试方法提升安全防护
Effective Red-Teaming of Policy-Adherent Agents
- 用会说服的多智能体模拟恶意用户,突破规则限制
- 新基准tau-break能更真实评估系统被操纵风险
- 现有防御措施效果有限,需更强算法保护
面向退款、取消等严格政策场景的任务导向型大模型代理日益普及。如何确保其始终遵守规则,在拒绝违规请求的同时保持自然交互,成为关键挑战。为此,我们提出一种新型威胁模型,聚焦于企图利用规则代理谋取私利的恶意用户。基于此,我们构建CRAFT——一个多智能体红队系统,采用政策感知的说服策略,在客户服务场景中有效攻破规则遵循型代理,优于传统越狱方法(如DAN提示、情感操控、胁迫)。在现有tau-bench基础上,我们引入tau-break这一互补基准,用于严格评估代理对操纵行为的鲁棒性。最后,我们测试了若干简单但有效的防御策略,结果表明这些措施虽有一定保护作用,但仍显不足,凸显亟需更强大的、以研究为基础的安全机制来抵御对抗攻击。
原文摘要 · Abstract (English)
Task-oriented LLM-based agents are increasingly used in domains with strict policies, such as refund eligibility or cancellation rules. The challenge lies in ensuring that the agent consistently adheres to these rules and policies, appropriately refusing any request that would violate them, while still maintaining a helpful and natural interaction. This calls for the development of tailored design and evaluation methodologies to ensure agent resilience against malicious user behavior. We propose a novel threat model that focuses on adversarial users aiming to exploit policy-adherent agents for personal benefit. To address this, we present CRAFT, a multi-agent red-teaming system that leverages policy-aware persuasive strategies to undermine a policy-adherent agent in a customer-service scenario, outperforming conventional jailbreak methods such as DAN prompts, emotional manipulation, and coercive. Building upon the existing tau-bench benchmark, we introduce tau-break, a complementary benchmark designed to rigorously assess the agent's robustness against manipulative user behavior. Finally, we evaluate several straightforward yet effective defense strategies. While these measures provide some protection, they fall short, highlighting the need for stronger, research-driven safeguards to protect policy-adherent agents from adversarial attacks
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。