不同用户提问方式影响大模型响应,存在群体不公平现象。
Prompt Fairness: Sub-group Disparities in LLMs
- 用信息论指标衡量提示词对模型输出的影响差异。
- 某些群体内部响应波动大,跨群体差异最高达0.28。
- 通过提示中性化和多轮生成可显著降低不公平性。
大型语言模型虽在诸多应用中表现高效,但其响应质量存在显著差异。本文研究提示词公平性问题:相同问题因用户表述风格不同,可能引发模型不同回应。我们提出基于信息论的度量方法,捕捉两个维度的偏差——子群体敏感性(组内响应变异)与跨群体一致性(组间响应变异)。实证分析显示,特定人口子群体既表现出更高的内部变异,又与其它群体存在更大偏离,揭示模型行为中的结构性不公。为此,我们提出实用干预策略,包括多次生成取多数投票及提示词中性化,有效提升响应稳定性并增强群体公平性。实验表明,未处理时跨群体差异最高达0.28,多数在0.14至0.22之间;应用中性化与多生成策略后,最大差距降至0.22,多数低于0.17,体现各群体输出更稳定一致。
原文摘要 · Abstract (English)
Large Language Models (LLMs), though shown to be effective in many applications, can vary significantly in their response quality. In this paper, we investigate this problem of prompt fairness: specifically, the phrasing of a prompt by different users/styles, despite the same question being asked in principle, may elicit different responses from an LLM. To quantify this disparity, we propose to use information-theoretic metrics that can capture two dimensions of bias: subgroup sensitivity, the variability of responses within a subgroup and cross group consistency, the variability of responses across subgroups. Our analysis reveals that certain subgroups exhibit both higher internal variability and greater divergence from others. Our empirical analysis reveals that certain demographic sub groups experience both higher internal variability and greater divergence from others, indicating structural inequities in model behavior. To mitigate these disparities, we propose practical interventions, including majority voting across multiple generations and prompt neutralization, which together improve response stability and enhance fairness across user populations. In the experiments, we observe clear prompt sensitivity disparities across demographic subgroups: before mitigation, cross-group divergence values reach 0.28 and typically fall in the from 0.14 to 0.22 range. After applying our neutralization and multi generation strategy, these divergences consistently decrease, with the largest gap reduced to 0.22 and many distances falling to 0.17 or below, indicating more stable and consistent outputs across subgroups.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。