arXiv:2604.06833cs.CRcs.LG2026-04

在边缘设备上通过数据净化实现小模型安全对齐,防止有毒数据污染。

FedDetox: Robust Federated SLM Alignment via On-Device Data Sanitization

论文配图:FedDetox: Robust Federated SLM Alignment via On-Device Data Sanitization
图 1 · 摘自论文原文
  • 边缘设备本地识别并替换危险内容为拒答模板,变毒为盾。
  • 保持模型安全水平接近中心化训练,且不损失通用能力。
  • 适合资源受限的移动端小语言模型安全训练场景。

随着高质量公开数据日益稀缺,联邦学习(FL)为利用宝贵的私有用户数据提供了重要途径,同时保障隐私。然而,真实客户端数据常包含有害或不安全信息,导致我们定义的无意数据投毒问题,严重损害全局模型的安全对齐效果。为此,我们提出FedDetox,一个专为资源受限边缘设备设计的小语言模型(SLMs)鲁棒对齐框架。首先,通过知识蒸馏将大型安全对齐教师模型中的先进安全对齐能力迁移至轻量级学生分类器,适用于资源受限的边缘设备。具体而言,在联邦学习中进行人类偏好对齐时,边缘客户端在源头识别不安全样本,并用拒答模板替换,有效将潜在毒物转化为积极的安全信号。实验表明,该方法在不牺牲通用性能的前提下,使模型安全性达到与集中式基线相当的水平。

原文摘要 · Abstract (English)

As high quality public data becomes scarce, Federated Learning (FL) provides a vital pathway to leverage valuable private user data while preserving privacy. However, real-world client data often contains toxic or unsafe information. This leads to a critical issue we define as unintended data poisoning, which can severely damage the safety alignment of global models during federated alignment. To address this, we propose FedDetox, a robust framework tailored for Small Language Models (SLMs) on resource-constrained edge devices. We first employ knowledge distillation to transfer sophisticated safety alignment capabilities from large scale safety aligned teacher models into light weight student classifiers suitable for resource constrained edge devices. Specifically, during federated learning for human preference alignment, the edge client identifies unsafe samples at the source and replaces them with refusal templates, effectively transforming potential poisons into positive safety signals. Experiments demonstrate that our approach preserves model safety at a level comparable to centralized baselines without compromising general utility.

联邦学习小模型安全对齐边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。