为大模型设计端到端安全防护框架,输入输出双管齐下提升可信度。
DeepKnown-Guard: A Proprietary Model-Based Safety Response Framework for AI Agents
- 基于四层分类的输入风险识别,精准区分安全与高危请求。
- 输出采用RAG+微调模型,确保回复可追溯且无幻觉,召回率达99.3%。
- 专用于高风险场景,适合金融、医疗等严苛领域的AI应用部署。
随着大语言模型(LLMs)的广泛应用,其安全问题日益突出,严重制约了在关键领域的可信部署。本文提出一种新型安全响应框架,从输入和输出两端系统性保障LLM安全。输入端采用监督微调的安全分类模型,基于细粒度四层分类体系(安全、不安全、有条件安全、关注焦点),实现对用户查询的精准风险识别与差异化处理,显著提升风险覆盖范围与业务场景适应性,风险召回率达99.3%。输出端融合检索增强生成(RAG)与专用微调解释模型,确保所有响应基于实时可信知识库,消除信息编造并支持结果溯源。实验表明,该安全控制模型在公开安全评估基准上得分显著优于基线模型TinyR1-Safety-8B;在自研高风险测试集上,各组件均达到100%安全得分,验证了其在复杂风险场景下的卓越防护能力。本研究为构建高安全性、高可信度的LLM应用提供了有效工程路径。
原文摘要 · Abstract (English)
With the widespread application of Large Language Models (LLMs), their associated security issues have become increasingly prominent, severely constraining their trustworthy deployment in critical domains. This paper proposes a novel safety response framework designed to systematically safeguard LLMs at both the input and output levels. At the input level, the framework employs a supervised fine-tuning-based safety classification model. Through a fine-grained four-tier taxonomy (Safe, Unsafe, Conditionally Safe, Focused Attention), it performs precise risk identification and differentiated handling of user queries, significantly enhancing risk coverage and business scenario adaptability, and achieving a risk recall rate of 99.3%. At the output level, the framework integrates Retrieval-Augmented Generation (RAG) with a specifically fine-tuned interpretation model, ensuring all responses are grounded in a real-time, trustworthy knowledge base. This approach eliminates information fabrication and enables result traceability. Experimental results demonstrate that our proposed safety control model achieves a significantly higher safety score on public safety evaluation benchmarks compared to the baseline model, TinyR1-Safety-8B. Furthermore, on our proprietary high-risk test set, the framework's components attained a perfect 100% safety score, validating their exceptional protective capabilities in complex risk scenarios. This research provides an effective engineering pathway for building high-security, high-trust LLM applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。