arXiv:2603.02588cs.CL2026-03被引 2

针对金融医疗等专业领域,打造抗攻击的AI内容安全防护模型。

ExpGuard: LLM Content Moderation in Specialized Domains

  • 专为金融、医疗、法律领域设计,识别技术性有害内容。
  • 在专家标注测试集上,对提示词和回复的分类准确率分别领先现有模型8.9%和15.3%。
  • 开源数据集与模型,支持跨领域安全防护研究。

随着大语言模型在现实应用中的普及,建立可靠的安全部署机制以监管其输入输出变得至关重要。当前的防护模型多聚焦通用人机交互,难以应对富含专业术语的特定领域中的有害及对抗性内容。为此,我们提出ExpGuard,一种专用于金融、医疗和法律领域的鲁棒防护模型。同时,我们构建了ExpGuardMix数据集,包含58,928条标注过的提示词及其对应的拒绝与合规响应,分为训练集ExpGuardTrain和由领域专家标注的高质量测试集ExpGuardTest。在ExpGuardTest及八个公开基准上的综合评估显示,ExpGuard整体表现优异,并在对抗性攻击下展现出显著韧性,提示词分类性能优于WildGuard达8.9%,回复分类性能提升15.3%。为促进研究发展,我们开源代码、数据与模型,支持向更多领域扩展,助力构建更强大的安全部署系统。

原文摘要 · Abstract (English)

With the growing deployment of large language models (LLMs) in real-world applications, establishing robust safety guardrails to moderate their inputs and outputs has become essential to ensure adherence to safety policies. Current guardrail models predominantly address general human-LLM interactions, rendering LLMs vulnerable to harmful and adversarial content within domain-specific contexts, particularly those rich in technical jargon and specialized concepts. To address this limitation, we introduce ExpGuard, a robust and specialized guardrail model designed to protect against harmful prompts and responses across financial, medical, and legal domains. In addition, we present ExpGuardMix, a meticulously curated dataset comprising 58,928 labeled prompts paired with corresponding refusal and compliant responses, from these specific sectors. This dataset is divided into two subsets: ExpGuardTrain, for model training, and ExpGuardTest, a high-quality test set annotated by domain experts to evaluate model robustness against technical and domain-specific content. Comprehensive evaluations conducted on ExpGuardTest and eight established public benchmarks reveal that ExpGuard delivers competitive performance across the board while demonstrating exceptional resilience to domain-specific adversarial attacks, surpassing state-of-the-art models such as WildGuard by up to 8.9% in prompt classification and 15.3% in response classification. To encourage further research and development, we open-source our code, data, and model, enabling adaptation to additional domains and supporting the creation of increasingly robust guardrail models.

内容安全大模型专业领域防护模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。