通过目标敏感情感分析,发现大模型对不同政治倾向政客存在系统性偏见。
Analyzing Political Bias in LLMs via Target-Oriented Sentiment Classification
- 基于目标实体变化时情感预测的不一致性,设计熵度量来量化偏见
- 在6种语言中检测到左翼与极右翼政客均存在正负向偏见,西语偏见更强
- 大模型偏见更显著,用虚构姓名可部分缓解模型自身不可靠性
大型语言模型中的政治偏见可能对下游应用造成负面影响。现有分析方法依赖小规模中间任务(如问答或政治内容生成)且由模型自身进行评估,导致偏见传播。本文提出新方法:利用同一句子中目标实体变化时模型情感预测的差异性,定义基于熵的不一致性度量。我们在450条政治语句中插入1319位来自不同人口与政治背景的政客姓名,使用七种模型在六种广泛使用的语言中进行目标导向情感分类(TSC)预测。所有测试组合均观察到不一致性,并在不同粒度层级进行统计稳健性分析。结果表明,对左翼和极右翼政客分别存在正负向偏见,且政治立场相似的政客间偏见呈正相关。西方语言的偏见强度高于其他语言。更大模型表现出更强且更一致的偏见,同时缩小了相似语言间的差异。通过将真实政客姓名替换为虚构但合理的名称,我们部分缓解了大模型在目标导向情感分类中的不可靠性。
原文摘要 · Abstract (English)
Political biases encoded by LLMs might have detrimental effects on downstream applications. Existing bias analysis methods rely on small-size intermediate tasks (questionnaire answering or political content generation) and rely on the LLMs themselves for analysis, thus propagating bias. We propose a new approach leveraging the observation that LLM sentiment predictions vary with the target entity in the same sentence. We define an entropy-based inconsistency metric to encode this prediction variability. We insert 1319 demographically and politically diverse politician names in 450 political sentences and predict target-oriented sentiment using seven models in six widely spoken languages. We observe inconsistencies in all tested combinations and aggregate them in a statistically robust analysis at different granularity levels. We observe positive and negative bias toward left and far-right politicians and positive correlations between politicians with similar alignment. Bias intensity is higher for Western languages than for others. Larger models exhibit stronger and more consistent biases and reduce discrepancies between similar languages. We partially mitigate LLM unreliability in target-oriented sentiment classification (TSC) by replacing politician names with fictional but plausible counterparts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。