提出多维度评测框架,揭示大模型偏见随提示设计剧烈波动
BiAxisBias: Evaluating LLM Bias Beyond a Single Prompt and a Single Explanation

- 设计跨任务/角色/视角/情感/措辞的200条刻板印象测试,系统评估偏见
- 17.1%的模型-陈述对在不同提示下改变选择,34.2%的响应存在输出矛盾
- 发现任务-情感交互是最大敏感因素,随机采样优于集中采样但分层效果有限
大模型偏见评分高度依赖审计设计。本文提出BiAxisBias,一个预定义的审计框架,通过在200条刻板印象陈述上系统变化任务、角色、视角、情感和措辞,同时保留强制选择与理由作为独立读出。其主矩阵涵盖8个大模型和401个模板(共641,600次响应)。在五个等价问题中,1,600个模型-陈述对中有17.1%改变选择;在三词表述子集上,不稳定性平均为10.5%,显著高于三组相同表述的6.3%。在四个受控任务范式中,28对模型中有9对反转结果;七模型因子敏感性分析显示,任务×情感是最大的二阶效应(原始eta-squared = 0.0465)。在10,000次等预算重抽样中,单模板集中采样的均方误差为10.63,全矩阵随机采样为1.86,条件分层采样为1.75;排名反转率分别为20.6%、4.8%和5.5%。因此,广泛覆盖带来主要增益,分层仅小幅降低误差,无排名优势。强制输出诊断显示,任务特定选择映射与人工编码理由立场在7,959个双重有效响应中存在34.2%分歧(31.2%与3.0%方向差异),表明输出契约敏感性,而非同一构念的双重验证。
原文摘要 · Abstract (English)
LLM bias scores can depend on audit design. We introduce BiAxisBias, a prespecified audit varying task, role, perspective, sentiment, and wording over 200 stereotype statements while retaining forced Selection and Rationale as separate protocol readouts. Its main matrix spans eight LLMs and 401 templates (641,600 responses). Across five equivalent questions, 17.1% of 1,600 model-statement pairs change Selection. With three observations per unit in both arms, instability averages 10.5% across all ten three-wording subsets, versus 6.3% for three identical calls. Across four controlled task paradigms, 9/28 model pairs reverse; a seven-model factorial sensitivity identifies task-by-sentiment as the largest two-way component (raw eta-squared = 0.0465). In 10,000 equal-budget resampling draws, mean absolute error against a declared 18-condition finite reference is 10.63 points for one-template concentration, 1.86 for matrix-wide simple random sampling, and 1.75 for condition stratification; ranking inversions are 20.6%, 4.8%, and 5.5%. Thus broad coverage drives the gain, while stratification has only a modest score-error advantage and no ranking advantage over random sampling. In a separate forced-output diagnostic, task-specific Selection mappings and judge-coded Rationale stance disagree in 34.2% of 7,959 dual-valid responses (31.2% versus 3.0% by direction). This diagnoses output-contract sensitivity, not two validated measures of one construct.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。