Protect让企业级大模型安全防护跨文本、图像、音频三模态统一运行。
Protect: Towards Robust Guardrailing Stack for Trustworthy Enterprise LLM Systems
- 用低秩适配微调多模态安全分类器,支持跨模态实时检测
- 在毒性、性别歧视等四个维度上超越WildGuard和GPT-4.1
- 适合金融、医疗等强监管领域部署,支持可审计安全管控
大型语言模型在企业及关键领域的广泛应用,凸显了构建鲁棒安全防护系统以保障安全、可靠与合规的迫切需求。现有方案普遍存在实时监控难、多模态处理弱、可解释性差等问题,难以在受监管环境中落地。传统防护机制多孤立运行于文本,无法适应多模态生产环境。本文提出Protect,一个原生支持文本、图像与音频输入的多模态安全防护模型,专为企业级部署设计。Protect通过在覆盖四类安全维度(毒性、性别歧视、数据隐私、提示注入)的多模态数据集上,利用低秩适配(LoRA)微调特定类别适配器实现。其教师辅助标注流程结合推理与解释轨迹,生成高保真、上下文感知的标签。实验表明,Protect在所有安全维度上均达到当前最佳性能,优于WildGuard、LlamaGuard-4及GPT-4.1等开源与闭源模型。Protect为可审计、可信赖、生产就绪的安全系统奠定了坚实基础,支持跨文本、图像、音频模态的统一防护。
原文摘要 · Abstract (English)
The increasing deployment of Large Language Models (LLMs) across enterprise and mission-critical domains has underscored the urgent need for robust guardrailing systems that ensure safety, reliability, and compliance. Existing solutions often struggle with real-time oversight, multi-modal data handling, and explainability -- limitations that hinder their adoption in regulated environments. Existing guardrails largely operate in isolation, focused on text alone making them inadequate for multi-modal, production-scale environments. We introduce Protect, natively multi-modal guardrailing model designed to operate seamlessly across text, image, and audio inputs, designed for enterprise-grade deployment. Protect integrates fine-tuned, category-specific adapters trained via Low-Rank Adaptation (LoRA) on an extensive, multi-modal dataset covering four safety dimensions: toxicity, sexism, data privacy, and prompt injection. Our teacher-assisted annotation pipeline leverages reasoning and explanation traces to generate high-fidelity, context-aware labels across modalities. Experimental results demonstrate state-of-the-art performance across all safety dimensions, surpassing existing open and proprietary models such as WildGuard, LlamaGuard-4, and GPT-4.1. Protect establishes a strong foundation for trustworthy, auditable, and production-ready safety systems capable of operating across text, image, and audio modalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。