arXiv:2510.27077cs.CL2025-10被引 7

通过对比蒸馏与抗噪训练,提升大模型对齐安全性与鲁棒性。

Contrastive Knowledge Transfer and Robust Optimization for Secure Alignment of Large Language Models

  • 用对比蒸馏将教师模型的知识边界迁移到学生模型。
  • 在噪声输入下仍保持稳定输出,显著提升抗干扰能力。
  • 适合关注大模型安全对齐与鲁棒训练的研究者。

本文针对大语言模型在安全对齐与鲁棒性方面的局限,提出一种结合对比蒸馏与抗噪训练的微调方法。该方法冻结主干模型,通过蒸馏将教师模型的知识边界传递给学生模型,从而提升语义一致性和对齐精度。同时,在训练中引入噪声扰动与鲁棒优化约束,确保模型在噪声和不确定输入下仍能保持稳定的预测输出。整体框架包含蒸馏损失、鲁棒性损失和正则化项,构成统一优化目标,平衡对齐能力与抗干扰性能。通过多角度实验验证,包括蒸馏权重敏感性、计算资源受限与混合精度环境下的稳定性分析,以及数据噪声和分布偏移的影响,结果表明该方法在知识迁移、鲁棒性及整体安全性上显著优于现有基线,在多个关键指标上表现最佳。本工作不仅丰富了参数高效微调的理论体系,也为构建更安全可信的对齐机制提供了新方案。

原文摘要 · Abstract (English)

This paper addresses the limitations of large-scale language models in safety alignment and robustness by proposing a fine-tuning method that combines contrastive distillation with noise-robust training. The method freezes the backbone model and transfers the knowledge boundaries of the teacher model to the student model through distillation, thereby improving semantic consistency and alignment accuracy. At the same time, noise perturbations and robust optimization constraints are introduced during training to ensure that the model maintains stable predictive outputs under noisy and uncertain inputs. The overall framework consists of distillation loss, robustness loss, and a regularization term, forming a unified optimization objective that balances alignment ability with resistance to interference. To systematically validate its effectiveness, the study designs experiments from multiple perspectives, including distillation weight sensitivity, stability analysis under computation budgets and mixed-precision environments, and the impact of data noise and distribution shifts on model performance. Results show that the method significantly outperforms existing baselines in knowledge transfer, robustness, and overall safety, achieving the best performance across several key metrics. This work not only enriches the theoretical system of parameter-efficient fine-tuning but also provides a new solution for building safer and more trustworthy alignment mechanisms.

模型对齐鲁棒训练知识蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。