arXiv:2608.03166cs.AI2026-08中稿 · and presented at A…

用多智能体对抗测试揭露角色扮演大模型的隐藏缺陷。

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation

论文配图:Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation
图 1 · 摘自论文原文
  • 设计三智能体系统,通过六种递进攻击策略持续施压目标模型。
  • 多策略测试使模型鲁棒性平均下降0.17–0.20分,暴露单策略无法发现的失效模式。
  • 自动化评分与人工高度一致,适合研究者用于评估和改进角色扮演AI安全。

角色扮演语言代理(RPLAs)正广泛应用于医疗协助、客户服务和教育等高风险场景,其在对抗压力下保持角色一致性、伦理约束与行为连贯性至关重要。现有评估方法依赖静态基准或孤立单轮提示,难以捕捉长期交互中累积的行为失效。本文提出一个模块化多智能体平台,通过结构化多轮对话对RPLAs进行对抗性压力测试。系统包含三个智能体:采用六种渐进式对抗策略的策略驱动提问者代理、待测的靶向代理,以及自动评分行为在角色忠实度、偏离度、伦理偏差和一致性维度的评判代理。在三种人物设定和三种LLM家族上的实验表明,多策略对抗测试揭示了单策略测试无法发现的失效模式,平均使整体鲁棒性评分下降0.17–0.20分。跨模型验证显示Llama-3.3-70B、GPT-4o-mini和Claude-3.5-Haiku均呈现一致退化趋势,其中权威挑战与情绪操控为最有效攻击策略。自动化评分与人类判断高度一致($r = 0.82$,Fleiss' $κ= 0.71$)。本工作以开源平台形式发布,支持AI安全研究与可复现的RPLA评测。尽管框架能系统发现失效模式,我们仍强调对抗测试方法潜在的伦理风险,倡导负责任使用以提升AI安全性。

原文摘要 · Abstract (English)

Role-Playing Language Agents (RPLAs) are increasingly deployed in high-stakes applications such as healthcare assistance, customer support, and education, where maintaining consistent personas, ethical constraints, and behavioral coherence under adversarial pressure is critical. Existing evaluation approaches rely on static benchmarks or isolated single-turn prompts that fail to capture cumulative behavioral failures emerging over extended interactions. We present a modular multi-agent platform for adversarially stress-testing RPLAs through structured, multi-turn dialogue. The system coordinates three agents: a strategy-driven Interrogator Agent that applies six progressive adversarial strategies, a Target Agent representing the RPLA under evaluation, and an automated Judging Agent that scores behavior across role fidelity, drift, ethical deviation, and consistency dimensions. Through experiments across three personas and three LLM families, we demonstrate that multi-strategy adversarial evaluation reveals failure modes invisible to single-strategy testing, reducing overall robustness scores by 0.17--0.20 points on average. Cross-model validation confirms consistent degradation patterns across Llama-3.3-70B, GPT-4o-mini, and Claude-3.5-Haiku, with Authority Challenge and Emotional Manipulation emerging as the most effective attack strategies. Automated judging achieves strong human alignment ($r = 0.82$, Fleiss' $κ= 0.71$). This work is released as an open-source platform to support AI safety and reproducible RPLA benchmarking. While the framework enables systematic discovery of failure modes, we acknowledge potential ethical risks associated with adversarial testing methodologies and emphasize responsible usage for improving AI safety.

角色扮演对抗测试多智能体AI安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。