用多层隐变量原型实现高效可定制的LLM内容审核
Efficient LLM Moderation with Multi-Layer Latent Prototypes
- 用多层中间表示的原型捕捉潜在风险特征
- 在多个基准上达到顶尖效果且几乎无计算开销
- 适合需要灵活安全策略的部署场景
尽管现代大模型在后训练阶段已对齐人类价值观,但在部署时仍需稳健的内容审核以防止有害输出。现有方法普遍存在性能与效率的权衡,且难以满足用户个性化需求。为此,我们提出多层原型审核器(MLPM),一种轻量级、高度可定制的输入审核工具。通过利用多层中间表示的原型,提升审核质量的同时保持高效率。该方法对生成流程添加可忽略的额外开销,可无缝适配任意模型。MLPM在多样化的审核基准上达到当前最优性能,并展现出对不同规模模型家族的强大可扩展性。此外,它能顺利集成至端到端审核流程中,与输出审核技术结合后进一步提升响应安全性。整体上,本工作为安全、稳健、高效的LLM部署提供了实用且灵活的解决方案。
原文摘要 · Abstract (English)
Although modern LLMs are aligned with human values during post-training, robust moderation remains essential to prevent harmful outputs at deployment time. Existing approaches suffer from performance-efficiency trade-offs and are difficult to customize to user-specific requirements. Motivated by this gap, we introduce Multi-Layer Prototype Moderator (MLPM), a lightweight and highly customizable input moderation tool. We propose leveraging prototypes of intermediate representations across multiple layers to improve moderation quality while maintaining high efficiency. By design, our method adds negligible overhead to the generation pipeline and can be seamlessly applied to any model. MLPM achieves state-of-the-art performance on diverse moderation benchmarks and demonstrates strong scalability across model families of various sizes. Moreover, we show that it integrates smoothly into end-to-end moderation pipelines and further improves response safety when combined with output moderation techniques. Overall, our work provides a practical and adaptable solution for safe, robust, and efficient LLM deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。