arXiv:2410.08442cs.LGcs.AI2024-10被引 3

JurEE用小模型集成提升AI对话安全,更准更快更省。

JurEE not Judges: safeguarding llm interactions with small, specialised Encoder Ensembles

  • 用多个专用编码器模型组成集成系统,协同判断风险。
  • 在多个基准上表现优于基线,准确率高、速度快、成本低。
  • 适合客服机器人等需严格内容审核的场景,可调风险阈值。

我们提出JurEE,一种由高效编码器模型组成的集成系统,用于增强基于大语言模型的AI-用户交互中的安全防护。与现有依赖大模型作裁判的方法不同,JurEE能对多种常见风险提供概率化风险评估,而不仅限于文本输出。该方法融合多元数据源,并采用渐进式合成数据生成技术(包括大模型辅助增强),提升模型鲁棒性与性能。我们构建了一个内部基准,涵盖OpenAI Moderation Dataset和ToxicChat等知名数据集,结果显示JurEE显著优于基线模型,在准确性、速度和成本效率方面均有优势。其模块化设计支持用户自定义风险阈值,适用于各类安全相关应用。各专用编码器协同决策的机制不仅提高预测精度,也增强可解释性,为大规模内容审核提供更高效、高性能且经济的替代方案。

原文摘要 · Abstract (English)

We introduce JurEE, an ensemble of efficient, encoder-only transformer models designed to strengthen safeguards in AI-User interactions within LLM-based systems. Unlike existing LLM-as-Judge methods, which often struggle with generalization across risk taxonomies and only provide textual outputs, JurEE offers probabilistic risk estimates across a wide range of prevalent risks. Our approach leverages diverse data sources and employs progressive synthetic data generation techniques, including LLM-assisted augmentation, to enhance model robustness and performance. We create an in-house benchmark comprising of other reputable benchmarks such as the OpenAI Moderation Dataset and ToxicChat, where we find JurEE significantly outperforms baseline models, demonstrating superior accuracy, speed, and cost-efficiency. This makes it particularly suitable for applications requiring stringent content moderation, such as customer-facing chatbots. The encoder-ensemble's modular design allows users to set tailored risk thresholds, enhancing its versatility across various safety-related applications. JurEE's collective decision-making process, where each specialized encoder model contributes to the final output, not only improves predictive accuracy but also enhances interpretability. This approach provides a more efficient, performant, and economical alternative to traditional LLMs for large-scale implementations requiring robust content moderation.

内容安全模型集成轻量级

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。