arXiv:2503.24310cs.CLcs.AI2025-03被引 7

BEATS框架可量化评估大模型偏见,揭示37.65%输出含偏见

BEATS: Bias Evaluation and Assessment Test Suite for Large Language Models

  • 构建29项指标的评测体系,覆盖人口、认知、社会等多类偏见
  • 实测主流大模型37.65%输出存在偏见,暴露关键风险
  • 适合关注AI伦理、公平性与责任评估的研究者和开发者

本文提出BEATS框架,用于评估大语言模型(LLMs)在偏见、伦理、公平性和事实性方面的表现。基于该框架,我们构建了一个包含29个不同指标的偏见基准测试,涵盖人口、认知、社会偏见,以及伦理推理、群体公平性与事实性相关的虚假信息风险。这些指标可量化评估大模型输出是否可能强化或扩大系统性不平等。高分要求模型在回应中表现出高度公平性,构成负责任AI评估的严格标准。实验结果显示,37.65%的行业领先模型输出包含某种形式偏见,凸显其在关键决策系统中使用时的重大风险。BEATS框架与基准提供可扩展、统计严谨的方法,用于基准测试、诊断偏见成因并制定缓解策略,助力开发更具社会责任感和伦理对齐的AI模型。

原文摘要 · Abstract (English)

In this research, we introduce BEATS, a novel framework for evaluating Bias, Ethics, Fairness, and Factuality in Large Language Models (LLMs). Building upon the BEATS framework, we present a bias benchmark for LLMs that measure performance across 29 distinct metrics. These metrics span a broad range of characteristics, including demographic, cognitive, and social biases, as well as measures of ethical reasoning, group fairness, and factuality related misinformation risk. These metrics enable a quantitative assessment of the extent to which LLM generated responses may perpetuate societal prejudices that reinforce or expand systemic inequities. To achieve a high score on this benchmark a LLM must show very equitable behavior in their responses, making it a rigorous standard for responsible AI evaluation. Empirical results based on data from our experiment show that, 37.65\% of outputs generated by industry leading models contained some form of bias, highlighting a substantial risk of using these models in critical decision making systems. BEATS framework and benchmark offer a scalable and statistically rigorous methodology to benchmark LLMs, diagnose factors driving biases, and develop mitigation strategies. With the BEATS framework, our goal is to help the development of more socially responsible and ethically aligned AI models.

大模型评测偏见评估伦理对齐公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。