arXiv:2509.00673cs.CLcs.AI2025-09ACL

对比了大模型在仇恨言论检测中的表现,发现删减版更准但更固执。

Confident, Calibrated, or Complicit: Safety Alignment and Ideological Bias in LLM Hate Speech Detection

  • 用政治人设测试模型,发现删减版比未删减版更抗意识形态影响。
  • 删减模型严格准确率达69.0%,未删减为64.1%,但都难识别讽刺语境。
  • 模型自述信心过高,对不同群体检测不公平,需改进评估框架。

我们研究了大语言模型在检测隐性与显性仇恨言论方面的有效性,比较了在部署场景下使用政治人设时,安全对齐程度较低(未删减)与较高(删减)模型的表现。尽管未删减模型常被认为更具开放性,但结果揭示出权衡:删减模型在准确率与鲁棒性上均优于未删减版本,严格准确率分别为69.0%与64.1%。然而,删减模型对人设影响的抵抗更强,而未删减模型更易受意识形态框架影响。此外,所有模型在理解讽刺等微妙语言时均出现严重失败。我们还发现性能在不同目标群体间存在显著公平差异,且普遍存在系统性过度自信,导致自我报告的置信度不可靠。这些发现挑战了大模型作为客观裁决者的观念,强调需要更复杂的审计框架,兼顾公平性、校准性和意识形态一致性。总体而言,应将‘部署中的删减’而非‘孤立的安全对齐’视为解释模型差异的更合适框架。

原文摘要 · Abstract (English)

We investigate the efficacy of Large Language Models (LLMs) in detecting implicit and explicit hate speech, examining how models with minimal safety alignment (uncensored) compare with more heavily aligned (censored) counterparts in a deployed-model setting when deployed using political personas. While uncensored models are often framed as offering a less constrained perspective, our results reveal a trade-off: censored models outperform their uncensored counterparts in both accuracy and robustness, achieving 69.0\% versus 64.1\% strict accuracy. However, this higher performance is also associated with greater resistance to persona-based influence, while uncensored models are more malleable to ideological framing. Furthermore, we identify critical failures across all models in understanding nuanced language such as irony. We also find alarming fairness disparities in performance across different targeted groups and systemic overconfidence that renders self-reported certainty unreliable. These findings challenge the notion of LLMs as objective arbiters and highlight the need for more sophisticated auditing frameworks that account for fairness, calibration, and ideological consistency. Taken together, these results point to censorship-as-deployed rather than safety alignment in isolation as the more appropriate frame for interpreting model differences.

仇恨言论检测模型偏差安全性公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。