arXiv:2412.09751cs.CYcs.AI2024-12被引 1

红队测试AI安全需关注价值观、劳动与心理影响

AI red-teaming is a sociotechnical problem: on values, labor, and harms

  • 跨学科合作研究红队工作的社会技术系统
  • 揭示红队工作背后的伦理假设与心理负担
  • 适合关注AI治理与人机协作的研究者阅读

随着生成式AI在现实场景中广泛应用,测试其性能与安全性愈发重要。红队已成为主流的AI模型测试方法,被企业优先采用,并写入政策法规。红队成员扮演攻击者角色,探测AI系统的安全机制并发现漏洞。然而,我们对这一工作及其影响知之甚少。本文呼吁计算机科学家与社会科学家协同研究红队相关的社会技术系统,避免重蹈内容审核工作的覆辙。文章强调需理解红队背后的价值观与假设、相关的劳动组织形式,以及对红队人员的心理影响,借鉴内容审核领域的经验教训。

原文摘要 · Abstract (English)

As generative AI technologies find more and more real-world applications, the importance of testing their performance and safety seems paramount. "Red-teaming" has quickly become the primary approach to test AI models--prioritized by AI companies, and enshrined in AI policy and regulation. Members of red teams act as adversaries, probing AI systems to test their safety mechanisms and uncover vulnerabilities. Yet we know far too little about this work or its implications. This essay calls for collaboration between computer scientists and social scientists to study the sociotechnical systems surrounding AI technologies, including the work of red-teaming, to avoid repeating the mistakes of the recent past. We highlight the importance of understanding the values and assumptions behind red-teaming, the labor arrangements involved, and the psychological impacts on red-teamers, drawing insights from the lessons learned around the work of content moderation.

AI治理红队测试人机协作社会技术

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。