在联邦学习中提升大模型安全性,防止有害内容传播。
Responsible Federated LLMs via Safety Filtering and Constitutional AI
- 将安全过滤与宪法AI引入联邦学习框架
- 在AdvBench上安全性能提升超20%
- 适合关注AI安全与可信部署的研究者
近期研究越来越多地采用联邦学习训练大语言模型(即FedLLM),但负责任人工智能(RAI)在该场景下仍缺乏探索。在FedLLM中,客户端训练数据可能包含有害内容,导致生成不安全响应。将此类模型聚合为全局模型并分发回客户端,会带来广泛部署不安全模型的风险。为此,我们引入两种成熟的RAI技术:安全过滤与宪法AI。实验表明,这些方法显著提升了模型安全性,在AdvBench上实现超过20%的性能提升。
原文摘要 · Abstract (English)
Recent research has increasingly focused on training large language models (LLMs) using federated learning, known as FedLLM. However, responsible AI (RAI), which aims to ensure safe and trustworthy responses, remains underexplored in this context. In FedLLM, client-side training data may contain harmful content, resulting in unsafe LLMs that can generate inappropriate responses. Aggregating such models into a global model and redistributing it to clients risks the widespread deployment of unsafe LLMs. To address this, we incorporate two well-established RAI techniques into FedLLM: safety filtering and constitutional AI. Our experiments show that these methods significantly improve LLM safety, achieving over 20% improvement on AdvBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。