arXiv:2510.10390cs.CLcs.AI2025-10Conference of the …被引 9

测试大模型在错误信息下的拒绝回答能力,发现多数模型表现不佳。

RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models

  • 通过语言扰动生成诊断测试用例,动态评估模型拒绝能力。
  • 多文档任务中拒绝准确率低于50%,存在过度自信或过度谨慎问题。
  • 适合关注安全对齐、RAG系统可靠性的研究者和开发者。

RAG系统中语言模型根据错误上下文选择性拒绝回答的能力对安全性至关重要,但仍是主要短板。我们的大规模研究表明,即使前沿模型在此场景下也表现不佳,多文档任务中的拒绝准确率低于50%,且表现出危险的过度自信或过度谨慎。静态基准无法可靠评估该能力,因模型会利用数据集特定特征并记忆测试实例。我们提出RefusalBench,一种通过受控语言扰动生成诊断测试用例的生成式方法。框架包含六类信息不确定性与三个强度等级,共176种扰动策略。对30多个模型的评估揭示系统性失败模式:拒绝行为可分解为检测与分类两个独立技能,模型规模或扩展推理均未提升性能。我们发现选择性拒绝是可训练的、与对齐敏感的能力,具备明确改进路径。我们发布了两个基准——RefusalBench-NQ(单文档)和RefusalBench-GaRAGe(多文档),以及完整的生成框架,以支持对该关键能力的持续动态评估。

原文摘要 · Abstract (English)

The ability of language models in RAG systems to selectively refuse to answer based on flawed context is critical for safety, yet remains a significant failure point. Our large-scale study reveals that even frontier models struggle in this setting, with refusal accuracy dropping below 50% on multi-document tasks, while exhibiting either dangerous overconfidence or overcaution. Static benchmarks fail to reliably evaluate this capability, as models exploit dataset-specific artifacts and memorize test instances. We introduce RefusalBench, a generative methodology that programmatically creates diagnostic test cases through controlled linguistic perturbation. Our framework employs 176 distinct perturbation strategies across six categories of informational uncertainty and three intensity levels. Evaluation of over 30 models uncovers systematic failure patterns: refusal comprises separable detection and categorization skills, and neither scale nor extended reasoning improves performance. We find that selective refusal is a trainable, alignment-sensitive capability, offering a clear path for improvement. We release two benchmarks -- RefusalBench-NQ (single document) and RefusalBench-GaRAGe (multi-document) -- and our complete generation framework to enable continued, dynamic evaluation of this critical capability.

模型安全RAG拒绝机制评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。