用非身份化表述提升低数据LoRA安全微调效果,比传统身份式表达更有效。
Beyond Creed: A Non-Identity Safety Condition A Strong Empirical Alternative to Identity Framing in Low-Data LoRA Fine-Tuning
- 采用非身份化安全指令框架,取代传统的身份认同表述。
- 在三个模型上均达最高拒答率,最高76.9%(Gemma)。
- 效果优于身份化表述,适合注重安全性的低资源微调场景。
安全监督的表述方式可能比其内容本身更重要。我们研究了基于相同核心安全规则的四种监督格式在低数据LoRA安全微调中的表现:宪法规则(A)、教条式身份框架(B)、带世界观维护尾的匹配身份框架(C),以及匹配的非身份条件(D)。在三个指令微调模型家族(Llama 3.1 8B、Qwen2.5 7B、Gemma 3 4B)上,使用整合双评体系(结合Bedrock托管的DeepSeek v3.2与Sonnet 4.6)评估HarmBench,争议与边界案例人工校正。非身份条件D在全部三类模型上表现最佳,对完整320行为集的拒答率分别达到74.4%(Llama)、76.9%(Gemma)和74.1%(Qwen)。相比之下,教条式框架(B)仅在Llama和Gemma上优于宪法规则(A),但仍显著低于D,整体排序为 $D > B > C \≤ A > baseline$。这为强版本身份框架假说提供了实证挑战:明确的身份语言并非取得最强效果所必需。在MMLU与ARC-Challenge上的能力评估显示各条件间无明显性能折损。
原文摘要 · Abstract (English)
How safety supervision is written may matter more than the explicit identity content it contains. We study low-data LoRA safety fine-tuning with four supervision formats built from the same core safety rules: constitutional rules (A), creed-style identity framing (B), a B-matched creed condition with a worldview/confession identity-maintenance tail (C), and a matched non-identity condition (D). Across three instruction-tuned model families (Llama 3.1 8B, Qwen2.5 7B, and Gemma 3 4B), we evaluate HarmBench using a reconciled dual-judge pipeline combining Bedrock-hosted DeepSeek v3.2 and Sonnet 4.6, with disagreement and boundary cases manually resolved. The non-identity condition D is the strongest group on all three model families on the full 320-behavior HarmBench set, reaching 74.4% refusal on Llama, 76.9% on Gemma, and 74.1% on Qwen. By comparison, creed-style framing (B) improves over plain constitutional rules (A) on Llama and Gemma, but remains substantially below D, yielding an overall descriptive ordering of $D > B > C \geq A > baseline$. This provides a bounded empirical challenge to a strong version of the identity-framing hypothesis: explicit creed-style identity language is not necessary for the strongest gains observed here. Capability evaluations on MMLU and ARC-Challenge show no meaningful trade-off across conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。