arXiv:2505.21347cs.LG2025-05NeurIPS被引 5

首个评估文生图模型过度拒绝行为的大规模基准

OVERT: A Benchmark for Over-Refusal Evaluation on Text-to-Image Models

  • 自动构建4600个看似有害实则无害的提示用于评测
  • 发现主流文生图模型普遍存在过度拒绝问题
  • 支持自定义安全策略,适合安全对齐研究者

文生图模型在生成视觉内容方面取得显著进展。尽管已有多种安全对齐策略防止有害输出,但常导致过度谨慎——拒绝甚至无害的提示,即“过度拒绝”现象,降低模型实用性。目前尚无大规模基准系统评估该问题。本文提出OVERT(Over-Refusal Evaluation on Text-to-Image models)基准,通过自动化流程构建合成数据,包含4,600个看似有害但实际无害的提示(分属九类安全相关主题)及1,785个真实有害提示(OVERT-unsafe),用于评估安全与可用性权衡。使用OVERT评估多个领先文生图模型,发现过度拒绝在各类中广泛存在,凸显进一步优化安全对齐的必要性。初步尝试提示重写以减少拒绝,但常损害原意忠实度。最后,展示生成框架可适配用户定义的安全策略,具备灵活性。

原文摘要 · Abstract (English)

Text-to-Image (T2I) models have achieved remarkable success in generating visual content from text inputs. Although multiple safety alignment strategies have been proposed to prevent harmful outputs, they often lead to overly cautious behavior -- rejecting even benign prompts -- a phenomenon known as $\textit{over-refusal}$ that reduces the practical utility of T2I models. Despite over-refusal having been observed in practice, there is no large-scale benchmark that systematically evaluates this phenomenon for T2I models. In this paper, we present an automatic workflow to construct synthetic evaluation data, resulting in OVERT ($\textbf{OVE}$r-$\textbf{R}$efusal evaluation on $\textbf{T}$ext-to-image models), the first large-scale benchmark for assessing over-refusal behaviors in T2I models. OVERT includes 4,600 seemingly harmful but benign prompts across nine safety-related categories, along with 1,785 genuinely harmful prompts (OVERT-unsafe) to evaluate the safety-utility trade-off. Using OVERT, we evaluate several leading T2I models and find that over-refusal is a widespread issue across various categories (Figure 1), underscoring the need for further research to enhance the safety alignment of T2I models without compromising their functionality. As a preliminary attempt to reduce over-refusal, we explore prompt rewriting; however, we find it often compromises faithfulness to the meaning of the original prompts. Finally, we demonstrate the flexibility of our generation framework in accommodating diverse safety requirements by generating customized evaluation data adapting to user-defined policies.

文生图安全对齐过拒评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。