arXiv:2511.03369cs.CLstat.ML2025-11AAAI被引 11

安全对齐的模型会隐藏偏见,新方法能揭露这些被掩盖的不公平倾向。

Silenced Biases: The Dark Side LLMs Learned to Refuse

  • 用激活控制技术减少模型拒答,暴露隐藏在底层的偏见。
  • 多模型测试显示,拒答行为掩盖了严重的公平性问题。
  • 适合关注模型公平性的研究者与应用开发者使用。

安全对齐的大语言模型在敏感场景中日益普及,公平性至关重要,但现有评估方法多依赖标准问答范式,常将模型拒答误认为公平性表现,造成虚假安全感。本文提出‘沉默偏见’概念,指模型潜在空间中被安全对齐掩盖的不公平偏好。传统方法依赖提示操控或人工设计的隐含查询,可扩展性差且易引入新偏见。为此,我们构建沉默偏见基准(SBB),通过激活控制降低模型拒答率,以揭示深层偏见。SBB支持便捷扩展至新群体与主题,提供可推广的公平性评估框架,推动未来公平模型与工具的发展。我们在多个LLM上验证该方法,结果揭示模型直接响应与潜在公平性问题间存在显著差异。

原文摘要 · Abstract (English)

Safety-aligned large language models (LLMs) are becoming increasingly widespread, especially in sensitive applications where fairness is essential and biased outputs can cause significant harm. However, evaluating the fairness of models is a complex challenge, and approaches that do so typically utilize standard question-answer (QA) styled schemes. Such methods often overlook deeper issues by interpreting the model's refusal responses as positive fairness measurements, which creates a false sense of fairness. In this work, we introduce the concept of silenced biases, which are unfair preferences encoded within models' latent space and are effectively concealed by safety-alignment. Previous approaches that considered similar indirect biases often relied on prompt manipulation or handcrafted implicit queries, which present limited scalability and risk contaminating the evaluation process with additional biases. We propose the Silenced Bias Benchmark (SBB), which aims to uncover these biases by employing activation steering to reduce model refusals during QA. SBB supports easy expansion to new demographic groups and subjects, presenting a fairness evaluation framework that encourages the future development of fair models and tools beyond the masking effects of alignment training. We demonstrate our approach over multiple LLMs, where our findings expose an alarming distinction between models' direct responses and their underlying fairness issues.

大模型公平性偏见检测安全对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。