对比了语言模型偏见评估中两种指标的优劣,发现没有哪个更胜一筹。
Textual Entailment is not a Better Bias Metric than Token Probability
- 用自然语言推理(NLI)替代词概率(TP)作为偏见度量方法
- 实验显示NLI与TP相关性低,且结果更不稳定
- 适合关注评估方法可靠性的研究人员参考
衡量语言模型中的社会偏见通常采用词概率(TP)指标,虽适用广泛但被批评为与真实使用场景和危害脱节。本文测试自然语言推理(NLI)作为替代偏见度量方法。在七个语言模型家族的大量实验中,我们发现NLI与TP在偏见评估上表现显著不同,不同NLI指标间相关性极低,且与TP指标相关性也很弱。NLI指标更脆弱、更不稳定,对反刻板印象句子的措辞变化稍不敏感,而对刻板印象表述的措辞变化稍更敏感。基于此矛盾证据,我们得出结论:在所有情况下,既不存在优于另一的偏见度量方法。尚无充分证据支持以NLI完全取代TP进行偏见评估。
原文摘要 · Abstract (English)
Measurement of social bias in language models is typically by token probability (TP) metrics, which are broadly applicable but have been criticized for their distance from real-world language model use cases and harms. In this work, we test natural language inference (NLI) as an alternative bias metric. In extensive experiments across seven LM families, we show that NLI and TP bias evaluation behave substantially differently, with very low correlation among different NLI metrics and between NLI and TP metrics. NLI metrics are more brittle and unstable, slightly less sensitive to wording of counterstereotypical sentences, and slightly more sensitive to wording of tested stereotypes than TP approaches. Given this conflicting evidence, we conclude that neither token probability nor natural language inference is a ``better'' bias metric in all cases. We do not find sufficient evidence to justify NLI as a complete replacement for TP metrics in bias evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。