arXiv:2605.05682cs.HCcs.AI2026-05

让AI模拟不同身份视角,更全面发现生成式AI的安全风险

PersonaTeaming: Supporting Persona-Driven Red-Teaming for Generative AI

论文配图:PersonaTeaming: Supporting Persona-Driven Red-Teaming for Generative AI
图 1 · 摘自论文原文
  • 用角色设定驱动提示词生成,拓展对抗性策略覆盖范围
  • 相比顶尖方法,攻击成功率更高且提示多样性保持良好
  • 支持用户自定义角色,促进人机协作激发创新思路

近期生成式AI安全研究强调红队测试需考虑测试者背景与视角对策略及风险发现的影响。现有自动化红队方法忽视人类身份特征,极少融合人工输入。本文提出PersonaTeaming工作流,将角色设定融入对抗性提示生成过程,探索更广泛攻击策略。相比当前最优的RainbowPlus方法,该工作流在保持提示多样性的同时实现更高攻击成功率。为弥补自动化角色对真实视角的近似局限,进一步构建PersonaTeaming Playground交互界面,使红队人员可自定义角色,并与AI协同演化和优化提示。11位行业从业者参与的用户研究表明,该平台催生多样化策略与有效输出,且AI建议虽未被完全采纳,仍能激发突破性思维。本工作推动了自动化与人机协同红队测试的发展,揭示了生成式AI红队中人机协作的关键设计启示。

原文摘要 · Abstract (English)

Recent developments in AI safety research have called for red-teaming methods that effectively surface potential risks posed by generative AI models, with growing emphasis on how red-teamers' backgrounds and perspectives shape their strategies and the risks they uncover. While automated red-teaming approaches promise to complement human red-teaming through larger-scale exploration, existing automated approaches do not account for human identities and rarely incorporate human inputs. In this work, we explore persona-driven red-teaming to advance both automated red-teaming and human-AI collaboration. We first develop PersonaTeaming Workflow, which incorporates personas into the adversarial prompt generation process to explore a wider spectrum of adversarial strategies. Compared to RainbowPlus, a state-of-the-art automated red-teaming method, PersonaTeaming Workflow achieves higher attack success rates while maintaining prompt diversity. However, since automated personas only approximate real human perspectives, we further instantiate PersonaTeaming Workflow as PersonaTeaming Playground, a user-facing interface that enables red-teamers to author their own personas and collaborate with AI to mutate and refine prompts. In a user study with 11 industry practitioners, we found that PersonaTeaming Playground enabled diverse red-teaming strategies and outputs that practitioners perceived as useful, and that AI-generated suggestions in the PersonaTeaming Playground encouraged out-of-the-box thinking even when practitioners did not follow them strictly. Together, our work advances both automated and human-in-the-loop approaches to red-teaming, while shedding light on interaction patterns and design insights for supporting human-AI collaboration in generative AI red-teaming.

红队测试人机协作安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。