构建16000条含结构化推理的有毒提示数据集,缓解大模型误拒安全请求问题。
FalseReject: A Resource for Improving Contextual Safety and Mitigating Over-Refusals in LLMs via Structured Reasoning
- 设计图神经网络增强的多智能体框架生成复杂提示,附带显式推理路径。
- 在29个主流大模型上验证,使用该数据集微调后误拒率显著下降。
- 适用于需要精准安全判断的高敏感场景,如医疗、法律等专业对话系统。
大型语言模型的安全对齐常导致对无害查询的过度拒绝,严重影响其在敏感场景中的可用性。为此,我们提出 FalseReject,一个包含16,000条看似有毒的查询及其在44个安全类别下的结构化回应的综合性资源。我们采用图感知的对抗性多智能体交互框架生成多样且复杂的提示,并通过显式推理结构帮助模型准确区分安全与不安全语境。FalseReject 包含针对标准指令微调模型和推理导向模型的训练数据,以及一个人工标注的基准测试集。我们在29个最先进的大模型上进行了广泛评测,结果揭示了持续存在的过度拒绝问题。实证表明,使用 FalseReject 进行监督微调能显著减少不必要的拒绝,同时不损害整体安全性或通用语言能力。
原文摘要 · Abstract (English)
Safety alignment approaches in large language models (LLMs) often lead to the over-refusal of benign queries, significantly diminishing their utility in sensitive scenarios. To address this challenge, we introduce FalseReject, a comprehensive resource containing 16k seemingly toxic queries accompanied by structured responses across 44 safety-related categories. We propose a graph-informed adversarial multi-agent interaction framework to generate diverse and complex prompts, while structuring responses with explicit reasoning to aid models in accurately distinguishing safe from unsafe contexts. FalseReject includes training datasets tailored for both standard instruction-tuned models and reasoning-oriented models, as well as a human-annotated benchmark test set. Our extensive benchmarking on 29 state-of-the-art (SOTA) LLMs reveals persistent over-refusal challenges. Empirical results demonstrate that supervised finetuning with FalseReject substantially reduces unnecessary refusals without compromising overall safety or general language capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。