反向宪法AI生成有毒数据,自动提升模型安全测试能力
Reverse Constitutional AI: A Framework for Controllable Toxic Data Generation via Probability-Clamped RLAIF

- 将无害宪法反转为毒性强宪法,通过迭代批改生成对抗数据
- 概率钳制使毒性数据语义连贯性提升15%且不削弱攻击性
- 适合做模型安全评估的自动化红队测试,无需人工标注
保障大语言模型的安全需依赖强大的红队测试,但高质量有毒数据的系统化生成仍鲜有研究。本文提出反向宪法AI(R-CAI),一种自动化、可控制的对抗性数据生成框架,突破了单一越狱提示的局限。通过将无害宪法反转为毒性强宪法,并借助批判-修订流水线迭代优化模型输出,R-CAI 实现了无需人工标注的多维度对抗数据规模化生成。然而,仅优化毒性奖励会导致奖励黑客行为和语义连贯性下降。为此,我们在基于AI反馈的强化学习中引入概率钳制,稳定对抗优化过程的同时保留对抗意图。实验表明,R-CAI 能生成多样且高质量的有毒数据,概率钳制使语义连贯性提升15%,且未牺牲对抗强度。整体上,R-CAI 提供了一个全自动化的红队数据生成与对齐模型系统性安全评估框架。
原文摘要 · Abstract (English)
Ensuring the safety of large language models (LLMs) requires robust red teaming, yet the systematic synthesis of high-quality toxic data remains under-explored. We propose Reverse Constitutional AI (R-CAI), a framework for automated and controllable adversarial data generation that moves beyond isolated jailbreak prompts. By inverting a harmless constitution into a constitution of toxicity and iteratively refining model outputs through a critique--revision pipeline, R-CAI enables scalable synthesis of multi-dimensional adversarial data without human annotation. Optimizing solely for toxicity-related rewards, however, can lead to reward hacking and degraded semantic coherence. To address this challenge, we introduce probability clamping within reinforcement learning from AI feedback, which stabilizes adversarial optimization while preserving adversarial intent. Experiments demonstrate that R-CAI generates diverse, high-quality toxic data and that probability clamping substantially improves semantic coherence (15%) without sacrificing adversarial strength. Overall, R-CAI provides a fully automated framework for red teaming data generation and systematic safety evaluation of aligned language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。