不同微调目标影响大模型安全与稳定性,规模越大越关键
Objective Matters: Fine-Tuning Objectives Shape Safety, Robustness, and Persona Drift
- 对比六种微调目标,固定其他变量研究其对安全的影响
- 大规模训练下,监督与偏好微调导致对抗脆弱性与人格漂移加剧
- 约束学习信号的方法(如ORPO、KL正则)显著提升鲁棒性与人格一致性
在良性数据上微调大语言模型仍可能导致对齐退化和对抗鲁棒性下降,但微调目标如何影响这些安全结果尚缺乏直接分析。本文在数据、领域、架构和优化过程保持一致的前提下,系统比较了六种微调目标:监督微调(SFT)、直接偏好优化(DPO)、条件微调(CFT)、免疫提示(Inoculation Prompting)、几率比偏好优化(ORPO)和KL正则化微调。在封闭形式推理与开放式生成任务中,发现目标选择会引发系统性、依赖训练规模的安全-能力权衡变化。小规模训练时,各目标的鲁棒性相近,但能力差异明显;大规模训练时,监督与偏好类方法使能力提升伴随显著的对抗脆弱性和人格漂移,而通过约束学习信号(尤其是ORPO和KL正则化)的方法能有效缓解上述问题。因此,微调目标在小规模时对安全性影响有限,但在大规模训练下成为决定对抗鲁棒性和潜在人格稳定性的关键因素。
原文摘要 · Abstract (English)
Fine-tuning LLMs on benign data can still degrade alignment and adversarial robustness, yet direct analysis of the role of fine-tuning objectives in shaping these safety outcomes remain limited. We present a controlled comparison of six fine-tuning objectives -- Supervised Fine-Tuning, Direct Preference Optimization, Conditional Fine-Tuning, Inoculation Prompting, Odds Ratio Preference Optimization, and KL-regularized fine-tuning -- holding data, domain, architecture, and optimization fixed. Across closed-form reasoning and open-ended generation tasks, we find that objective choice induces systematic, scale-dependent shifts along the safety-capability frontier. At small training budgets, robustness is similar across objectives but capability differs. At larger budgets, objectives diverge sharply: supervised and preference-based tuning tightly couple capability gains to increased adversarial vulnerability and persona drift, while objectives that constrain learning signals -- especially ORPO and KL-regularization -- substantially mitigate both. Fine-tuning objectives therefore matter little for safety at small scales but become a primary driver of adversarial robustness and latent persona stability as training scale increases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。