首个可定制公平性校准的综合偏见评测框架,能精准识别语言模型对各国的系统性偏见。
SAGED: A Holistic Bias-Benchmarking Pipeline for Language Models with Customisable Fairness Calibration
- 构建五阶段全流程评测管道,覆盖数据采集到偏见诊断,支持自定义公平性基准。
- 在20国语料上测试8B级模型,发现所有模型均对俄、中等国存在显著偏见,尤其针对中国。
- 提出反事实分支与基线校准机制,有效缓解提示词和评估工具带来的偏见干扰。
有偏大型语言模型的开发被广泛认为至关重要,但现有评测基准因范围有限、数据污染及缺乏公平性基线而难以检测偏见。SAGED(bias)是首个综合性评测管道,解决上述问题。该管道包含五个核心阶段:数据抓取、基准构建、响应生成、数值特征提取与偏见诊断,采用最大差异度量(如影响比)和偏见集中度量(如最大Z分数)。针对评估工具偏见和提示上下文偏差可能扭曲结果的问题,引入反事实分支与基线校准以缓解。以全球G20国家为对象,测试Gemma2、Llama3.1、Mistral和Qwen2等主流8B级模型。情感分析显示,Mistral与Qwen2的最高偏见差异较低、偏见集中度更高,但所有模型均明显偏向俄罗斯和(除Qwen2外)中国。进一步实验中,模型扮演美国总统角色时,偏见加剧且方向不一;其中Qwen2与Mistral几乎不参与角色扮演,而Llama3.1与Gemma2对特朗普的角色扮演远超拜登与哈里斯,揭示模型在角色扮演任务中的性能偏见。
原文摘要 · Abstract (English)
The development of unbiased large language models is widely recognized as crucial, yet existing benchmarks fall short in detecting biases due to limited scope, contamination, and lack of a fairness baseline. SAGED(bias) is the first holistic benchmarking pipeline to address these problems. The pipeline encompasses five core stages: scraping materials, assembling benchmarks, generating responses, extracting numeric features, and diagnosing with disparity metrics. SAGED includes metrics for max disparity, such as impact ratio, and bias concentration, such as Max Z-scores. Noticing that metric tool bias and contextual bias in prompts can distort evaluation, SAGED implements counterfactual branching and baseline calibration for mitigation. For demonstration, we use SAGED on G20 Countries with popular 8b-level models including Gemma2, Llama3.1, Mistral, and Qwen2. With sentiment analysis, we find that while Mistral and Qwen2 show lower max disparity and higher bias concentration than Gemma2 and Llama3.1, all models are notably biased against countries like Russia and (except for Qwen2) China. With further experiments to have models role-playing U.S. presidents, we see bias amplifies and shifts in heterogeneous directions. Moreover, we see Qwen2 and Mistral not engage in role-playing, while Llama3.1 and Gemma2 role-play Trump notably more intensively than Biden and Harris, indicating role-playing performance bias in these models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。