arXiv:2510.09004cs.CL2025-10

用低成本方法实现大模型安全对齐而不降性能

Decoupling Safety into Orthogonal Subspace: Cost-Efficient and Performance-Preserving Alignment for Large Language Models

  • 通过LoRA微调仅安全数据,实现零性能损失的安全对齐
  • 实验证明安全能力可被解耦到与主模型空间正交的低秩子空间
  • 适合追求高效安全增强的AI研发团队使用

安全对齐对构建可信人工智能至关重要,但现有方法需高成本搜索安全与通用能力的平衡比例,收益有限。本文发现基于LoRA的拒绝训练即使仅在安全数据上训练,也能保持模型性能的同时实现安全对齐,表明LoRA是低成本、保性能且即插即用的安全补丁。我们从理论和实验两方面证明,LoRA能将安全能力解耦至与模型内在变换空间近似正交的低秩子空间,确保安全增强不干扰原有能力。

原文摘要 · Abstract (English)

Safety alignment is essential for building trustworthy artificial intelligence, yet it remains challenging to enhance model safety without degrading general performance. Current approaches require computationally expensive searches for the optimal proportion of safety-critical and general-purpose data to balance safety and general performance, incurring high costs with limited gains. In this work, we show that LoRA-based Refusal-training enables performance-preserving safety alignment even when trained solely on safety data, demonstrating that LoRA serves as cost-efficient, performance-preserving, and plug-and-play safety patches. Beyond empirical findings, we provide both theoretical and experimental evidence that LoRA effectively decouples safety into a low-rank subspace largely orthogonal to the model's intrinsic transformation space, ensuring that safety enhancements do not interfere with inherent capabilities.

安全对齐LoRA大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。