arXiv:2510.00232cs.CLcs.AI2025-10中稿 · ICLR被引 3

构建统一基准评估大模型去偏方法,更贴近真实使用场景。

BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses

  • 将现有数据重构为问答式响应格式,统一评估环境。
  • 提出响应级评分指标,衡量输出公平性与安全性。
  • 对比8种主流方法,分析提示与训练范式的优劣。

现有大模型去偏方法采用不同基线和评估指标,导致比较不一致。且多数评估仅基于模型对有偏与无偏上下文的概率差异,忽视用户实际阅读响应时对公平、安全输出的期待。为此,我们提出 BiasFreeBench,一个实证基准,通过将现有数据集重新组织为统一的查询-响应格式,在多选题问答与开放多轮问答两种场景下,系统比较8种主流去偏技术(4种基于提示,4种基于训练)。我们引入响应级指标 Bias-Free Score,衡量模型输出在公平性、安全性及反刻板印象方面的表现。在提示与训练范式、模型规模、不同训练策略对未见偏见类型的泛化能力等维度上进行系统分析。我们公开该基准,旨在建立统一的去偏研究测试平台。

原文摘要 · Abstract (English)

Existing studies on bias mitigation methods for large language models (LLMs) use diverse baselines and metrics to evaluate debiasing performance, leading to inconsistent comparisons among them. Moreover, their evaluations are mostly based on the comparison between LLMs' probabilities of biased and unbiased contexts, which ignores the gap between such evaluations and real-world use cases where users interact with LLMs by reading model responses and expect fair and safe outputs rather than LLMs' probabilities. To enable consistent evaluation across debiasing methods and bridge this gap, we introduce BiasFreeBench, an empirical benchmark that comprehensively compares eight mainstream bias mitigation techniques (covering four prompting-based and four training-based methods) on two test scenarios (multi-choice QA and open-ended multi-turn QA) by reorganizing existing datasets into a unified query-response setting. We further introduce a response-level metric, Bias-Free Score, to measure the extent to which LLM responses are fair, safe, and anti-stereotypical. Debiasing performances are systematically compared and analyzed across key dimensions: the prompting vs. training paradigm, model size, and generalization of different training strategies to unseen bias types. We release our benchmark, aiming to establish a unified testbed for bias mitigation research.

大模型去偏评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。