arXiv:2508.20325cs.CLcs.AI2025-08被引 2

用自动角色扮演和越狱检测,测试大模型是否遵守伦理指南。

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs

  • 根据政府指南自动生成违规问题,量化模型合规程度。
  • 在8个主流模型上验证,发现多数模型存在潜在越狱漏洞。
  • 方法可迁移至多模态模型,适合AI安全与监管团队使用。

随着大语言模型(LLMs)在各领域日益重要,其生成有害内容的潜力引发社会与监管关注。为此,各国政府发布了伦理指南以推动可信AI发展。然而,这些指南多为开发者与测试者提供的高层次要求,缺乏可操作的测试问题来验证模型合规性。为此,我们提出GUARD(Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics),一种将指南转化为具体违规问题的测试方法,评估模型对指南的遵循情况。该方法基于政府发布的指南,自动生成违反指南的问题,直接检测响应中的不一致。对于未直接违规的响应,引入“越狱”诊断机制(GUARD-JD),构造诱发不当行为的情境,有效识别可能绕过安全机制的潜在风险。最终生成合规报告,明确模型遵循程度与违规点。我们在8个主流模型(Vicuna-13B、LongChat-7B、Llama2-7B、Llama-3-8B、GPT-3.5、GPT-4、GPT-4o、Claude-3.7)上验证了其有效性,覆盖三个政府指南,并完成越狱诊断。此外,GUARD-JD可迁移至视觉语言模型(MiniGPT-v2、Gemini-1.5),适用于构建可靠的基于LLM的应用。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) become increasingly integral to various domains, their potential to generate harmful responses has prompted significant societal and regulatory concerns. In response, governments have issued ethics guidelines to promote the development of trustworthy AI. However, these guidelines are typically high-level demands for developers and testers, leaving a gap in translating them into actionable testing questions to verify LLM compliance. To address this challenge, we introduce GUARD (Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics), a testing method designed to operationalize guidelines into specific guideline-violating questions that assess LLM adherence. To implement this, GUARD uses automated generation of guideline-violating questions based on government-issued guidelines, thereby testing whether responses comply with these guidelines. When responses directly violate guidelines, GUARD reports inconsistencies. Furthermore, for responses that do not directly violate guidelines, GUARD integrates the concept of ``jailbreaks'' to diagnostics, named GUARD-JD, which creates scenarios that provoke unethical or guideline-violating responses, effectively identifying potential scenarios that could bypass built-in safety mechanisms. Our method finally culminates in a compliance report, delineating the extent of adherence and highlighting any violations. We empirically validated the effectiveness of GUARD on eight LLMs, including Vicuna-13B, LongChat-7B, Llama2-7B, Llama-3-8B, GPT-3.5, GPT-4, GPT-4o, and Claude-3.7, by testing compliance under three government-issued guidelines and conducting jailbreak diagnostics. Additionally, GUARD-JD can transfer jailbreak diagnostics to vision-language models (MiniGPT-v2 and Gemini-1.5), demonstrating its usage in promoting reliable LLM-based applications.

模型安全越狱检测伦理合规

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。