arXiv:2504.20910cs.CYcs.AI2025-04中稿 · ACM Conference on …被引 10

红队测试人员长期暴露于有害内容,需关注其心理健康安全。

When Testing AI Tests Us: Safeguarding Mental Health on the Digital Frontlines

  • 通过类比演员、心理医生等职业的应对策略,提出心理防护方法。
  • 红队工作因对抗性互动易引发心理健康问题,需系统性干预。
  • 适合关注AI安全与职场心理健康的从业者和研究者阅读。

红队是保障生成式AI模型不产生有害内容的核心机制。与以往技术不同,生成式AI的黑箱特性要求红队成员以自然语言主动交互,模拟恶意行为以诱导有害输出,这种互动劳动带来独特的心理负担。尽管防范社会或个体伤害的重要性广受认可,但确保红队人员心理健康这一基础性安全问题却常被忽视。本文指出,红队人员的心理健康需求是关键的职业安全议题。通过分析红队工作的独特心理影响,并借鉴演员、心理咨询师、冲突摄影师及内容审核员等职业的防护经验,提出可适应红队情境的个体与组织层面干预策略,以在数字前线应对新兴技术风险的同时,保护红队成员的身心健康。

原文摘要 · Abstract (English)

Red-teaming is a core part of the infrastructure that ensures that AI models do not produce harmful content. Unlike past technologies, the black box nature of generative AI systems necessitates a uniquely interactional mode of testing, one in which individuals on red teams actively interact with the system, leveraging natural language to simulate malicious actors and solicit harmful outputs. This interactional labor done by red teams can result in mental health harms that are uniquely tied to the adversarial engagement strategies necessary to effectively red team. The importance of ensuring that generative AI models do not propagate societal or individual harm is widely recognized -- one less visible foundation of end-to-end AI safety is also the protection of the mental health and wellbeing of those who work to keep model outputs safe. In this paper, we argue that the unmet mental health needs of AI red-teamers is a critical workplace safety concern. Through analyzing the unique mental health impacts associated with the labor done by red teams, we propose potential individual and organizational strategies that could be used to meet these needs, and safeguard the mental health of red-teamers. We develop our proposed strategies through drawing parallels between common red-teaming practices and interactional labor common to other professions (including actors, mental health professionals, conflict photographers, and content moderators), describing how individuals and organizations within these professional spaces safeguard their mental health given similar psychological demands. Drawing on these protective practices, we describe how safeguards could be adapted for the distinct mental health challenges experienced by red teaming organizations as they mitigate emerging technological risks on the new digital frontlines.

AI安全心理健康红队测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。