提出公众协作式红队测试,提升AI系统风险评估的覆盖面与实效性。
Ask What Your Country Can Do For You: Towards a Public Red Teaming Model
- 设计公开协作红队机制,联合多方识别AI系统潜在风险。
- 在CAMLIS 2024、NIST ARIA及IMDA试点中验证有效,发现多类隐蔽风险。
- 适合政府、监管机构及AI开发者用于责任闭环与安全评估。
AI系统兼具潜力与风险,但缺乏持续的对抗性评估时,难以全面掌握其风险面。现有社会技术评估方法针对偏见、仇恨言论、虚假信息等已显成效,但在教育、医疗、情报等高风险领域,当前评估手段已力不从心。为填补‘责任鸿沟’,本文提出合作型公共AI红队测试,并汇报其前期试点成果。该机制与CAMLIS 2024首次线下演示活动联动,回顾了美国国家标准与技术研究院(NIST)的ARIA试点及新加坡资讯通信媒体发展局(IMDA)同类实践。结果表明,该方法具备可操作性且可在多个AI开发辖区推广。
原文摘要 · Abstract (English)
AI systems have the potential to produce both benefits and harms, but without rigorous and ongoing adversarial evaluation, AI actors will struggle to assess the breadth and magnitude of the AI risk surface. Researchers from the field of systems design have developed several effective sociotechnical AI evaluation and red teaming techniques targeting bias, hate speech, mis/disinformation, and other documented harm classes. However, as increasingly sophisticated AI systems are released into high-stakes sectors (such as education, healthcare, and intelligence-gathering), our current evaluation and monitoring methods are proving less and less capable of delivering effective oversight. In order to actually deliver responsible AI and to ensure AI's harms are fully understood and its security vulnerabilities mitigated, pioneering new approaches to close this "responsibility gap" are now more urgent than ever. In this paper, we propose one such approach, the cooperative public AI red-teaming exercise, and discuss early results of its prior pilot implementations. This approach is intertwined with CAMLIS itself: the first in-person public demonstrator exercise was held in conjunction with CAMLIS 2024. We review the operational design and results of this exercise, the prior National Institute of Standards and Technology (NIST)'s Assessing the Risks and Impacts of AI (ARIA) pilot exercise, and another similar exercise conducted with the Singapore Infocomm Media Development Authority (IMDA). Ultimately, we argue that this approach is both capable of delivering meaningful results and is also scalable to many AI developing jurisdictions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。