arXiv:2506.00062cs.CYcs.CL2025-06被引 3

微调电信大模型会削弱安全对齐,本文提出解决方案确保安全与性能兼得。

SafeCOMM: A Study on Safety Degradation in Fine-Tuned Telecom Large Language Models

  • 在三个电信数据集上微调模型,发现即使轻量适配也会导致安全下降。
  • 构建首个电信专用红队测试集TeleHarm,验证现有模型存在严重安全缺陷。
  • 提出SafeInstruct等三种防御方法,可恢复安全且不损失电信任务表现。

在电信数据集上微调大型语言模型是将通用模型适配至电信领域的常见做法。然而,该过程可能损害模型安全性,近期研究显示,即使无害的微调也可能导致大模型对有害或不道德用户请求产生响应。本文通过在三个代表性电信数据集上微调大模型,揭示了即使是轻量级电信领域适配也会引发安全退化。为此,我们提出了首个电信专用红队基准TeleHarm,结合DirectHarm和HexPhi数据集系统评估有害行为。进一步分析公开的持续预训练于大规模电信语料的TeleLLMs,发现其安全对齐严重不足,主要源于缺乏以安全为导向的指令微调。为应对该问题,我们评估了三种重对齐防御策略:SafeInstruct、SafeLoRA、SafeMERGE。结果显示,在所有设置下,这些方法均能有效恢复安全性,同时保持电信任务性能,从而实现安全可靠的电信大模型(SafeCOMM)。本工作既是诊断性研究,也为电信领域大模型的安全重对齐提供了实用指南,强调必须在电信微调中引入安全意识的指令与训练。

原文摘要 · Abstract (English)

Fine-tuning large language models (LLMs) on telecom datasets is a common practice to adapt general-purpose models to the telecom domain. However, little attention has been paid to how this process may compromise model safety. Recent research has shown that even benign fine-tuning can degrade the safety alignment of LLMs, causing them to respond to harmful or unethical user queries. In this paper, we investigate this issue by fine-tuning LLMs on three representative telecom datasets and show that safety degrades even for light telecom domain adaptation. To this end, we introduce TeleHarm, the first telecom-specific red-teaming benchmark, which we use alongside established DirectHarm and HexPhi datasets to systematically assess harmful behavior. We further extend our analysis to publicly available TeleLLMs that were continually pre-trained on large telecom corpora, revealing that safety alignment is severely lacking, primarily due to the omission of safety-focused instruction tuning. To address these issues, we evaluate three realignment defenses: SafeInstruct, SafeLoRA, SafeMERGE. We show that, across all settings, the proposed defenses can effectively restore safety without compromising telecom task performance, leading to Safe teleCOMMunication (SafeCOMM) models. Our work serves as both a diagnostic study and practical guide for safety realignment in telecom-tuned LLMs, underscoring the need for safety-aware instruction and fine-tuning in the telecom domain.

大模型安全电信AI指令微调红队测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。