提出四维评估标准,检验微调防御有效性是否真实可靠。
Acceptance Cards:A Four-Diagnostic Standard for Safe Fine-Tuning Defense Claims
- 用四维诊断法验证微调防御效果:统计可靠性、语义泛化、机制对齐、跨任务迁移。
- 重新评估SafeLoRA发现其在Gemma-2-2B-it上四项诊断中至少三项不达标。
- 适合关注模型安全评估严谨性的研究者和实践者参考。
安全微调防御常基于保留集上的差距缩小来背书,但该缩小可能源于采样噪声、主体偏差、能力退化或非可迁移机制。本文提出接受卡片(Acceptance Cards):一种评估协议、文档对象、可执行审计包及针对安全微调防御声明的证据标准。该协议在认定差距缩小为有效前,需通过统计可靠性、新语义泛化、机制对齐与跨任务迁移四项检验。在安装差距协议下重新评分,SafeLoRA在Gemma-2-2B-it上未通过全卡验证:严格机制分类编码下四项全败,宽松收缩重标码下仍失败三项。此为单一模型家族的窄范围审计,非对SafeLoRA整体有效性的全面判断。46个单元审计中,无一满足严格合取条件。最接近的家族仅勉强通过可靠性和机制检查,但未达新主体阈值,缺乏严格迁移通过,且存在可测量的部署精度损失。
原文摘要 · Abstract (English)
Safe fine-tuning defenses are often endorsed on the basis of a held-out gap reduction, but the same reduction can come from sampling noise, subject artifacts, capability loss, or a mechanism that does not transfer. We introduce Acceptance Cards: an evaluation protocol, a documentation object, an executable audit package, and a claim-specific evidential standard for safe fine-tuning defense claims. The protocol checks statistical reliability, fresh semantic generalization, mechanism alignment, and cross-task transfer before treating a gap reduction as a full-card pass. Re-scored under this installed-gap protocol, SafeLoRA fails the full-card pass on Gemma-2-2B-it: under strict mechanism-class coding it fails all four diagnostics, and under a permissive shrinkage relabel it still fails three of four. This is a narrow installed-gap audit on one model family, not a global judgment of SafeLoRA's effectiveness. In a 46-cell audit, no cell satisfies the strict conjunction. The closest family is a near miss that passes reliability and mechanism checks where the required data are available, but fails the fresh-subject threshold, lacks a strict transfer pass, and carries a measurable deployment-accuracy cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。