arXiv:2510.19005cs.CL2025-10EMNLP被引 1

动态构建测试集,精准捕捉大模型过度防御问题。

Dynamic Evaluation for Oversensitivity in LLMs

  • 基于模型特性动态生成挑战性数据集
  • 覆盖25个模型,共45万条样本的评测基准
  • 适合评估模型安全策略与交互体验的科研人员

过度敏感指大语言模型对本无害的提示产生防御性拒绝,干扰用户交互并模糊有害与无害内容的界限。现有评测依赖静态数据集,随模型演进而过时,导致数据污染和评估能力下降。为此,我们提出一种框架,动态生成针对特定模型的挑战数据,捕捉新兴防御模式,并与各模型行为对齐。基于此,我们构建OVERBENCH,涵盖25个模型家族的45万条样本,提供持续演进的过度敏感评估视角,可追踪模型发展中的防御触发点,揭示静态数据集忽略的漏洞。

原文摘要 · Abstract (English)

Oversensitivity occurs when language models defensively reject prompts that are actually benign. This behavior not only disrupts user interactions but also obscures the boundary between harmful and harmless content. Existing benchmarks rely on static datasets that degrade overtime as models evolve, leading to data contamination and diminished evaluative power. To address this, we develop a framework that dynamically generates model-specific challenging datasets, capturing emerging defensive patterns and aligning with each model's unique behavior. Building on this approach, we construct OVERBENCH, a benchmark that aggregates these datasets across diverse LLM families, encompassing 450,000 samples from 25 models. OVERBENCH provides a dynamic and evolving perspective on oversensitivity, allowing for continuous monitoring of defensive triggers as models advance, highlighting vulnerabilities that static datasets overlook.

大模型安全评测基准过度防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。