Granite Guardian 能识别多种风险,让大模型使用更安全。
Granite Guardian
- 用人工标注与合成数据训练,覆盖社会偏见、越狱攻击等多类风险
- 在有害内容和RAG幻觉检测上分别取得0.871和0.854的AUC分数
- 开源模型,适合关注AI安全与负责任开发的研究者和开发者
我们提出Granite Guardian系列防护模型,旨在为任意大语言模型(LLM)提供提示与回复的风险检测能力,实现安全可控的应用。该模型全面覆盖社会偏见、粗俗语、暴力、性内容、不道德行为、越狱攻击及检索增强生成(RAG)中的幻觉风险,如上下文相关性、事实一致性与回答相关性。基于来自多个来源的人工标注数据与合成数据进行训练,解决了传统风险检测模型忽略越狱攻击与RAG特定问题的局限。在有害内容与RAG幻觉相关基准测试中,分别获得0.871与0.854的AUC分数,是当前最通用且最具竞争力的模型。作为开源项目发布,旨在推动社区负责任的AI发展。
原文摘要 · Abstract (English)
We introduce the Granite Guardian models, a suite of safeguards designed to provide risk detection for prompts and responses, enabling safe and responsible use in combination with any large language model (LLM). These models offer comprehensive coverage across multiple risk dimensions, including social bias, profanity, violence, sexual content, unethical behavior, jailbreaking, and hallucination-related risks such as context relevance, groundedness, and answer relevance for retrieval-augmented generation (RAG). Trained on a unique dataset combining human annotations from diverse sources and synthetic data, Granite Guardian models address risks typically overlooked by traditional risk detection models, such as jailbreaks and RAG-specific issues. With AUC scores of 0.871 and 0.854 on harmful content and RAG-hallucination-related benchmarks respectively, Granite Guardian is the most generalizable and competitive model available in the space. Released as open-source, Granite Guardian aims to promote responsible AI development across the community. https://github.com/ibm-granite/granite-guardian
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。