小模型公平性评估新框架,揭示看似中立实则脆弱的隐藏风险
Beyond Bias Scores: Unmasking Vacuous Neutrality in Small Language Models
- 提出多维度评估框架VaNeu,覆盖偏见、效用、模糊处理等四阶段
- 9个0.5-5B参数小模型测试显示,早期低偏见模型后续仍存漏洞
- 适用于部署前审查,尤其适合社会敏感场景的负责任应用
小型语言模型(SLMs)在资源受限场景中的快速采用,已超出对其伦理与公平性的理解。为此,我们提出空洞中立性框架(VaNeu),一种用于部署前评估SLM公平性的多维范式。该框架从四个阶段考察模型鲁棒性:偏见、效用、模糊处理能力及位置偏见,涵盖多样社会偏见类别。据我们所知,这是首次对0.5-5B参数范围内的小模型进行大规模审计,填补了介于BERT类编码器与主流大模型之间的“中间层级”空白。我们在模糊与明确上下文中评估了来自四个模型家族的九个广泛使用的小模型。结果显示,早期表现出低偏见的模型在后续评估中常出现失败,暴露出隐藏缺陷与不可靠推理。这些发现强调了全面理解小模型公平性与可靠性的重要性,并将该框架定位为社会敏感场景中负责任部署的可靠工具。
原文摘要 · Abstract (English)
The rapid adoption of Small Language Models (SLMs) for resource constrained applications has outpaced our understanding of their ethical and fairness implications. To address this gap, we introduce the Vacuous Neutrality Framework (VaNeu), a multi-dimensional evaluation paradigm designed to assess SLM fairness prior to deployment. The framework examines model robustness across four stages - biases, utility, ambiguity handling, and positional bias over diverse social bias categories. To the best of our knowledge, this work presents the first large-scale audit of SLMs in the 0.5-5B parameter range, an overlooked "middle tier" between BERT-class encoders and flagship LLMs. We evaluate nine widely used SLMs spanning four model families under both ambiguous and disambiguated contexts. Our findings show that models demonstrating low bias in early stages often fail subsequent evaluations, revealing hidden vulnerabilities and unreliable reasoning. These results underscore the need for a more comprehensive understanding of fairness and reliability in SLMs, and position the proposed framework as a principled tool for responsible deployment in socially sensitive settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。