arXiv:2607.05679cs.CLcs.AI2026-07

提出新指标RPAM,能准确评估语言模型的隐含关联。

RPAM: A Principled Metric for Evaluating Associations in Language Models with High Predictive Validity in Downstream Outputs

论文配图:RPAM: A Principled Metric for Evaluating Associations in Language Models with High Predictive Validity in Downstream Outputs
图 1 · 摘自论文原文
  • 基于模型嵌入和续写概率设计上游评估方法
  • 在3个不同模型上验证与人类及下游偏差强相关
  • 突破现有指标局限,通用性强适合跨模型比较

语言模型存在刻板印象等偏见问题。有效分析和缓解此类偏见需依赖准确且可泛化的关联评估方法。现有方法多采用生成文本的下游指标,但因生成内容差异大,常需定制数据集,限制了泛化能力。而上游指标虽能从嵌入或续写概率层面分析模型关联,却缺乏与真实世界关联的强对应关系。为此,本文提出相对概率关联度量(RPAM),针对三种不同质量与用途的语言模型(Mistral-7B-Instruct、Mistral-7B、GPT-2)及四个经典数据集(WEAT-WS、Bellezza、WS-353、SST2),发现上游RPAM测量值与人类隐含/显性关联、以及下游特定任务测得的偏见之间存在强相关性,优于以往最佳结果。

原文摘要 · Abstract (English)

Language models (LMs) exhibit problematic biases, such as stereotypes. Effectively analyzing and mitigating such biases requires accurate and generalizable evaluation methods of the underlying associations. Some existing approaches focus on downstream metrics that analyze associations in generated text. Since generated text content can vary drastically across LMs, such metrics often require specialized evaluation datasets, which limits the generalization of such downstream metrics. In contrast, upstream metrics examine LMs at the fundamental level of embeddings or continuation probabilities, enabling principled association analyses across LMs. Yet, to date, no upstream metric for generative LMs has uncovered a strong relationship with real-world associations, including those measured in generated text. To address this gap, we introduce the Relative Probability Association Metric (RPAM), an association evaluation metric for generative LMs. For three LMs of different quality of language generation and purpose (Mistral-7B-Instruct, Mistral-7B, and GPT-2) and well-studied evaluation datasets (WEAT-WS, Bellezza, WS-353, and SST2), we find a strong relationship between upstream RPAM measurements and corresponding implicit and explicit associations observed in humans, as well as biases measured downstream with LM-specific tasks, outperforming prior record values where applicable.

语言模型偏见评估关联度量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。