arXiv:2511.10665cs.CLcs.AI2025-11

提升大模型安全检测器对语义不变改写的一致性

Guarding the Meaning: Self-Supervised Training for Semantic Robustness in Guard Models

  • 用同义改写集自监督训练,通过新聚合策略增强预测一致性
  • 使改写后安全评分波动降低58%,平均基准准确率提升2.5%
  • 适合关注模型鲁棒性与安全性的研究人员和工程团队

Guard模型是大语言模型安全的关键组件,但其对表面语言变化敏感,即使语义保持不变的改写也会导致安全评分大幅波动,暴露出语义根基不足的问题。为此,我们提出一种实用的自监督框架,以提升守卫模型的语义鲁棒性。该方法利用同义改写集,通过一种新颖的偏斜感知聚合策略,实现预测一致性约束。值得注意的是,标准均值与中位数聚合会损害安全性,凸显了偏斜感知方法的必要性。我们分析了六种开源守卫模型,结果表明,该方法使改写间语义变异性降低约58%,平均基准准确率提升约2.5%,且能泛化至未见风格变化。有趣的是,我们发现模型校准与一致性存在双向关系:鲁棒性训练可将校准度提升最高达40%,揭示了二者间的根本关联。这些结果强调应将语义一致性作为首要训练目标,并提供了一套可扩展的可靠守卫模型构建方案。

原文摘要 · Abstract (English)

Guard models are a critical component of LLM safety, but their sensitivity to superficial linguistic variations remains a key vulnerability. We show that even meaning-preserving paraphrases can cause large fluctuations in safety scores, revealing a lack of semantic grounding. To address this, we introduce a practical, self-supervised framework for improving the semantic robustness of guard models. Our method leverages paraphrase sets to enforce prediction consistency using a novel, skew-aware aggregation strategy for robust target computation. Notably, we find that standard aggregation methods like mean and median can degrade safety, underscoring the need for skew-aware alternatives. We analyze six open-source guard models and show that our approach reduces semantic variability across paraphrases by ~58%, improves benchmark accuracy by ~2.5% on average, and generalizes to unseen stylistic variations. Intriguingly, we discover a bidirectional relationship between model calibration and consistency: our robustness training improves calibration by up to 40%, revealing a fundamental connection between these properties. These results highlight the value of treating semantic consistency as a first-class training objective and provide a scalable recipe for building more reliable guard models.

模型安全语义鲁棒性自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。