arXiv:2604.16606cs.CRcs.LG2026-04

SafeLM统一提升大模型隐私、安全、防幻觉与抗攻击能力,适合高风险场景部署。

SafeLM: Unified Privacy-Aware Optimization for Trustworthy Federated Large Language Models

  • 融合联邦学习与加密梯度,兼顾隐私与通信效率。
  • 检测有害内容准确率达98.0%,梯度反演性能下降至15.1 dB。
  • 组件独立有效,集成后实现隐私与实用性的良好平衡。

大语言模型日益应用于高风险领域,但对其多重安全挑战的统一处理仍不足。本文提出SafeLM框架,协同应对隐私、安全、虚假信息与对抗鲁棒性四大问题:结合联邦训练与梯度智能优化及Paillier加密保障隐私;集成训练与推理阶段防御机制;采用对比对齐与校准解码减少幻觉;引入对齐感知二值聚合增强鲁棒性并保持有限重建质量。在事实性、毒性与成员推断等基准测试中,SafeLM实现98.0%有害内容检测准确率,通信量降低96.9%,梯度反演峰值信噪比从31.7 dB降至15.1 dB。消融实验表明各模块独立贡献,集成后达成隐私-效用权衡的最优表现,支持可信大型语言模型部署。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed in high-stakes domains, yet a unified treatment of their overlapping safety challenges remains lacking. We present SafeLM, a framework that jointly addresses four pillars of LLM safety: privacy, security, misinformation, and adversarial robustness. SafeLM combines federated training with gradient smartification and Paillier encryption for privacy, integrates defenses against training and inference-time attacks, employs contrastive grounding with calibrated decoding to reduce hallucinations, and introduces alignment-aware binarized aggregation to enhance robustness while maintaining bounded reconstruction quality. Across benchmarks on factuality, toxicity, and membership inference, SafeLM achieves 98.0% harmful content detection accuracy, reduces communication by 96.9%, and lowers gradient inversion PSNR from 31.7 dB to 15.1 dB. Ablations show that each component contributes independently, whereas their integration yields a strong privacy utility efficiency trade-off for deploying trustworthy LLMs.

联邦学习大模型安全隐私保护对抗鲁棒

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。