测试大模型安全防线的漏洞探测方法,提升对抗攻击的防御能力。
Global Challenge for Safe and Secure LLMs Track 1
- 设计自动化攻击技术,主动挖掘大模型的安全漏洞。
- 在多种违规场景中成功绕过内容过滤机制,验证防护短板。
- 适合关注AI安全、模型防御的研究者与从业者参考。
本文介绍了由新加坡人工智能(AISG)与网络安全研发计划办公室(CRPO)联合发起的全球大语言模型安全与可信挑战赛,旨在推动针对自动化越狱攻击的先进防御机制发展。随着大模型在医疗、金融及公共管理等关键领域的广泛应用,抵御对抗性攻击以防止滥用并维护伦理标准至关重要。该赛事设立两个赛道,其中第一赛道要求参赛者开发自动化方法,通过诱导生成不当响应来探测大模型漏洞,全面检验现有安全协议的可靠性。任务涵盖从辱骂性语言到虚假信息乃至非法活动等多种违规场景,旨在深入理解大模型的脆弱性,并为构建更稳健的模型提供依据。
原文摘要 · Abstract (English)
This paper introduces the Global Challenge for Safe and Secure Large Language Models (LLMs), a pioneering initiative organized by AI Singapore (AISG) and the CyberSG R&D Programme Office (CRPO) to foster the development of advanced defense mechanisms against automated jailbreaking attacks. With the increasing integration of LLMs in critical sectors such as healthcare, finance, and public administration, ensuring these models are resilient to adversarial attacks is vital for preventing misuse and upholding ethical standards. This competition focused on two distinct tracks designed to evaluate and enhance the robustness of LLM security frameworks. Track 1 tasked participants with developing automated methods to probe LLM vulnerabilities by eliciting undesirable responses, effectively testing the limits of existing safety protocols within LLMs. Participants were challenged to devise techniques that could bypass content safeguards across a diverse array of scenarios, from offensive language to misinformation and illegal activities. Through this process, Track 1 aimed to deepen the understanding of LLM vulnerabilities and provide insights for creating more resilient models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。