arXiv:2606.09866cs.LGcs.AI2026-06

让大模型在微调中既保持安全又不丢任务能力

Two to Tango: Coupled Task-Reference Selection for Safe LLM Fine-tuning

论文配图:Two to Tango: Coupled Task-Reference Selection for Safe LLM Fine-tuning
图 1 · 摘自论文原文
  • 联合选择安全参考样本和兼容任务数据,动态更新安全方向
  • 在10亿到80亿参数模型上,安全得分比最强基线高至少5.10分
  • 适合需要持续学习且注重安全性的实际部署场景

在下游数据上微调对齐安全的大语言模型(LLMs)虽能提升适配性,但可能损害已学得的安全行为。现有方法依赖固定安全样例、全局约束或单向任务过滤。我们诊断发现任务更新会暴露不同安全约束,由此提出双选机制(DualSelect),联合选择相关参考与兼容任务样本。该框架在任务更新前刷新条件化安全参考,再筛选与参考方向一致的任务样本。基于极小极大视角,通过熵正则化评分代理、懒惰参考刷新和梯度修正,同时选择高保真损失与任务冲突的参考,以及兼容样本。在1B-8B规模的LLM上,DualSelect在不损失任务效用的前提下保留了安全性能;使用REDORCA评估,其安全平均得分较最强基线至少提升5.10点,且在多数评测者下保持领先,仅带来适度开销。该思路可拓展至以保留为重点的持续学习。

原文摘要 · Abstract (English)

Fine-tuning safety aligned large language models (LLMs) on downstream data improves adaptation but may erode learned safety behavior. Existing methods use fixed safety examples, global constraints, or one-sided task filtering. Our diagnostics show task updates expose different safety constraints, motivating joint selection of relevant references and compatible task samples. We propose DualSelect, a coupled framework for task and reference selection that refreshes task conditioned safety references before filtering whole task samples compatible with the induced reference direction. Under a minimax view, DualSelect selects safety references with high preservation loss and task conflict, together with compatible task samples, through entropy-regularized scoring surrogates, lazy reference refresh, and gradient correction. On 1B-8B LLMs, DualSelect preserves safety without losing task utility; using the REDORCA judge, it improves Safety Avg. over the strongest baseline by at least 5.10 points and remains highest in Safety Avg. across judges with moderate overhead. This view extends to retention focused continual learning.

大模型安全微调持续学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。