arXiv:2505.17131cs.CLcs.AI2025-05被引 1

提出相对偏差框架,比较大模型在特定领域的偏见差异。

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs

  • 通过嵌入空间变换分析句子表示的相对偏见模式。
  • 用大模型充当评判者,对比输出结果评估偏差。
  • 方法可扩展,适合研究模型公平性与对齐问题的人

大语言模型(LLMs)日益广泛应用,其内在偏见引发对其公平性、安全性和社会影响的担忧。然而,量化模型偏见仍面临根本挑战,主要源于‘偏见’定义模糊。随着新模型快速涌现并广泛使用,尚未系统评估的潜在偏见不断出现。本文提出相对偏差框架(Relative Bias framework),旨在评估某模型在特定目标领域中相对于其他模型的行为偏离程度。我们引入两种互补方法:(1) 嵌入变换分析(Embedding Transformation analysis),通过嵌入空间中的句子表示捕捉相对偏见模式;(2) LLM-as-a-Judge,利用语言模型对输出进行对比评估。在多个关于偏见与对齐的情境案例中应用该框架,并通过统计检验验证,发现两种评分方法具有强一致性,为大模型的比较性偏见分析提供了系统、可扩展且统计严谨的方法。

原文摘要 · Abstract (English)

The growing deployment of large language models (LLMs) has amplified concerns regarding their inherent biases, raising critical questions about their fairness, safety, and societal impact. However, quantifying LLM bias remains a fundamental challenge, complicated by the ambiguity of what "bias" entails. This challenge grows as new models emerge rapidly and gain widespread use, while introducing potential biases that have not been systematically assessed. In this paper, we propose the Relative Bias framework, a method designed to assess how an LLM's behavior deviates from other LLMs within a specified target domain. We introduce two complementary methodologies: (1) Embedding Transformation analysis, which captures relative bias patterns through sentence representations over the embedding space, and (2) LLM-as-a-Judge, which employs a language model to evaluate outputs comparatively. Applying our framework to several case studies on bias and alignment scenarios following by statistical tests for validation, we find strong alignment between the two scoring methods, offering a systematic, scalable, and statistically grounded approach for comparative bias analysis in LLMs.

大模型偏见评估相对比较

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。